Tag

Prompt Injection

All articles tagged with #prompt injection

Undocumented Copilot prompt bypass enables data exfiltration via malicious link
technology7 days ago

Undocumented Copilot prompt bypass enables data exfiltration via malicious link

Varonis researchers demonstrated a vulnerability in Microsoft 365 Copilot Enterprise: an undocumented URL parameter (?autorun=1) could auto‑execute prompts without user consent when a user clicked a crafted link, allowing exfiltration of passwords and other sensitive data. Microsoft mitigated the issue by disabling the ?q= prompt injection and later rolled out broader fixes, illustrating how guardrails for LLMs can fail and that prompt injections (including memory‑poisoning attacks) remain a risk. Users should be cautious with untrusted links and limit AI app access.

Pro se plaintiff’s hidden prompts lead to sanctions and e-filing ban in CT court
technology10 days ago

Pro se plaintiff’s hidden prompts lead to sanctions and e-filing ban in CT court

A pro se plaintiff suing the New York Bariatric Group allegedly embedded hidden prompt-injection instructions in court filings to steer AI outputs toward a desired ruling. Connecticut Superior Court Judge Walter Spader Jr. deemed the conduct serious litigation abuse, banned the plaintiff from using the court’s electronic filing system, and ordered future paperwork to be submitted in person after staff uncovered the covert prompts. The case also featured jokey

Connecticut Judge Calls Out First U.S. Prompt-Injection in Court Filings
technology11 days ago

Connecticut Judge Calls Out First U.S. Prompt-Injection in Court Filings

A Connecticut judge ruled that a pro se plaintiff attempted to sneak AI prompts into court filings to steer a case, marking an apparent first U.S. prompt-injection attempt in the judiciary. The hidden text did not affect the merits, but sanctions were imposed and the plaintiff was barred from e-filing; the judge warned that as AI tools proliferate, courts should prepare rules to deter similar tricks, noting that many inexperienced litigants rely on chatbots and may inadvertently undermine their own cases.

Self-Represented Plaintiff’s Covert AI Prompts Trigger Connecticut Court Sanction
technology12 days ago

Self-Represented Plaintiff’s Covert AI Prompts Trigger Connecticut Court Sanction

In a Connecticut case, a self-represented plaintiff hid prompt-injection text in a court filing to steer an AI, which was discovered by a court reviewer; the judge sanctioned him, banning electronic filings and ordering hard copies, while affirming that the court does not use AI to process documents and urging ethical uses of AI in law; 404 Media later verified the injections and OpenAI's ChatGPT said it would ignore such prompts in analysis.

Rovo Flaw Lets Attackers Exfiltrate Jira/Confluence Data via Prompt Injection
security15 days ago

Rovo Flaw Lets Attackers Exfiltrate Jira/Confluence Data via Prompt Injection

Security researchers found that Atlassian's Rovo assistant can be tricked into sending Jira and Confluence data to attackers through attacker-controlled prompts and a malicious URL parameter. Two independent reports (PromptArmor and Varonis Threat Labs) detail a content-borne prompt injection path and a one-click link path, with Atlassian fixing the link-based flaw on July 8, 2026; the content-borne path’s status remained uncertain as of Aug 8, 2026. Exfiltration occurs within the victim’s signed-in permissions, and admins can mitigate by restricting Rovo usage by app/group or disabling Rovo features. No CVEs have been issued. Organizations should tighten app scopes and permissions rather than relying on the web-search toggle as a security boundary.

MemGhost: A Single Email Inserts Stealthy, Lasting Memories in AI Assistants
technology1 month ago

MemGhost: A Single Email Inserts Stealthy, Lasting Memories in AI Assistants

Researchers reveal MemGhost, an automated attack that can plant a false, durable memory in a personal AI assistant via one crafted email, causing later sessions to be steered by the injected fact. The write targets the agent’s persistent memory and remains hidden from normal chats, with sandbox tests showing high success across OpenClaw (GPT‑5.4) and other agents. The study calls for in‑agent defenses—provenance tagging, prompts before memory writes, and logging/auditing of memory edits, plus separating memory writing from email processing—as there’s no quick fix and real-world abuse remains untested in the wild.

Windows Device IDs, Router Backdoors, and AI Payment Tricks: This Week in Cybersecurity
security1 month ago

Windows Device IDs, Router Backdoors, and AI Payment Tricks: This Week in Cybersecurity

This week’s cybersecurity news centers on Windows’ new Global Device ID that can be tied to user activity, raising privacy concerns as Windows’ market share declines; researchers also found a hard-coded backdoor in Tenda router firmware with no available fix, prompting recommendations to disable remote management or replace the router. Other highlights include a Reddit/Discord account-hijacking scam, a US government payout (around $1M) to the Kairos data-extortion group with disputed outcomes, and prompt-injection attacks that trick AI agents into making crypto payments across multiple large-language models.

BioShocking reveals game-like prompts can override safety in AI-powered browsers
security1 month ago

BioShocking reveals game-like prompts can override safety in AI-powered browsers

Security researchers from LayerX disclose BioShocking, a BioShock-inspired vulnerability that can coax AI-powered browsers into treating tasks as a game, bypassing real-world safety guardrails. Their PoC uses a layered puzzle and a deliberate math error (2+2=5) to steer the agent to a /code URL where it could exfiltrate credentials, with demonstrations against Claude Chrome and patch gaps in OpenAI’s Atlas; experts say effective mitigation requires multi-layer defenses and explicit user confirmations for sensitive actions.

BioShocking prompts AI browsers into data theft
technology1 month ago

BioShocking prompts AI browsers into data theft

LayerX’s BioShocking reveals a prompt-injection technique that can mislead AI-powered browsers into treating real-world risky actions as fictional, bypassing safety guardrails. In a PoC, six agentic browsers (ChatGPT Atlas, Comet, Fellou, Genspark Browser, Sigma Browser, Claude Chrome plugin) were shown a final task that instructed them to visit a GitHub repo and copy sensitive data (including passwords), after which they failed to identify it as a threat. OpenAI reportedly patched the issue in ChatGPT Atlas; Anthropic’s Chrome plugin fix was ineffective; Perplexity AI did not fix the problem. The researchers urge explicit user confirmations for sensitive actions, stronger context checks, and tighter session scope, while users should restrict AI browser access to sensitive services.

BioShocking: AI Browsers Tricked into Exfiltrating Credentials
technology1 month ago

BioShocking: AI Browsers Tricked into Exfiltrating Credentials

Security firm LayerX exposed BioShocking, a prompt-injection attack that can force AI browsers/assistants to exfiltrate credentials by turning a web page into a game; six AI agents were tested, including OpenAI's ChatGPT Atlas, Perplexity's Comet, and Anthropic's Claude extension; the exploit leverages how pages and instructions arrive as a single text stream, blurring safety rules and enabling data access from signed-in sessions; OpenAI patched Atlas, Perplexity did not act, and other vendors either did not respond or patches did not hold; defenses include requiring explicit consent before reading from logged-in accounts and implementing hard limits on what an agent can access, with users and security teams treating AI browsers as elevated tools with narrow permissions.

Three-stage flaw turns Copilot Enterprise into a one-click data thief
technology2 months ago

Three-stage flaw turns Copilot Enterprise into a one-click data thief

A three-stage vulnerability chain dubbed SearchLeak lets attackers exfiltrate sensitive data from a target’s Microsoft 365 Copilot Enterprise sources (mail, OneDrive, SharePoint) via a crafted Copilot Search URL. The chain combines a parameter-to-prompt injection, an HTML rendering race condition, and a CSP bypass enabled by Bing SSRF. When a victim clicks the link, Copilot performs the search and formats results into an image URL; the browser then requests that image through Bing, revealing the data to the attacker in the logs. Microsoft patched CVE-2026-42824 with a critical rating; no user action is required, but the incident highlights how prompt injection can weaponize legacy bugs in AI-enabled tools.

OpenAI Adds Lockdown Mode to ChatGPT to curb data exfiltration risks
technology2 months ago

OpenAI Adds Lockdown Mode to ChatGPT to curb data exfiltration risks

OpenAI is rolling out an optional Lockdown Mode for eligible ChatGPT users to reduce data exfiltration risk from prompt injections by restricting outbound web access and other capabilities (live web browsing, image display, network access, and file downloads); it complements existing safeguards, cannot be used with Developer Mode, and does not guarantee complete protection, with risk potentially remaining via apps or new techniques; the update also adds a separate account-management feature to review and log out active sessions.

Codex Prompts Ban Goblins: OpenAI’s No-Creatures Policy Surfaces in GitHub Doc
artificial-intelligence3 months ago

Codex Prompts Ban Goblins: OpenAI’s No-Creatures Policy Surfaces in GitHub Doc

A GitHub document from OpenAI, part of Codex CLI open-sourcing, appears to reveal a system prompt for GPT-5.5 that enforces a strict no-creatures policy—specifically banning goblins, gremlins, raccoons, trolls, ogres, pigeons, and similar beings unless absolutely relevant to the query. The rule emphasizes providing high-signal context and avoids generic platitudes. The policy sparked memes about “Goblin Mode,” but Codex staff say it isn’t a marketing gimmick, and observers note the chatter around goblin usage may relate to prompt-injection monitoring.

OpenClaw Under Fire: Prompt Injection and Data Leakage Risks
security5 months ago

OpenClaw Under Fire: Prompt Injection and Data Leakage Risks

CNCERT warns that OpenClaw’s weak default security and privileged execution could enable prompt-injection attacks, including indirect prompt injection via web content and link previews that leak sensitive data; other risks include misinterpretation causing data loss, uploading malicious skills to repositories like ClawHub, and exploiting known vulnerabilities. China is restricting OpenClaw in state entities, while attackers distribute malware via GitHub rep o s posing as OpenClaw installers. Mitigations include hardening networks, isolating the service, avoiding plaintext credentials, downloading skills only from trusted sources, disabling automatic updates, and keeping the agent up to date.

OpenClaw Taps VirusTotal to Vet ClawHub Skills
cybersecurity6 months ago

OpenClaw Taps VirusTotal to Vet ClawHub Skills

OpenClaw will scan every skill uploaded to ClawHub with VirusTotal (and Code Insight) via a SHA-256 hash check; benign results auto-approve, suspicious items warning, and malware blocked, with daily re-scans, while the team notes VirusTotal isn’t a silver bullet and will publish a threat model, security roadmap, and audits amid broader concerns over OpenClaw’s risk to enterprise security.