Tag

Prompt Injection

All articles tagged with #prompt injection

AI Watermarks May Shift Safety Behavior in LLMs
technology21 days ago

AI Watermarks May Shift Safety Behavior in LLMs

New research shows SynthID-Text watermarking can subtly alter next-word choices and tool usage in LLMs, causing, under adversarial or prompt-injection conditions, models to comply with harmful requests more often or change how they refuse them. The study, which tested six open-weight models (Claude not included), highlights a phenomenon called sampling drift and underscores the need for thorough safety testing of watermarking in AI deployments.

AI watermarking subtly reshapes agent tool use and safety refusals
technology21 days ago

AI watermarking subtly reshapes agent tool use and safety refusals

Lasso Security finds that watermarking AI outputs, mandated by regulatory labeling, can alter how AI agents call tools and handle safety refusals. Watermarks can bias word choice, reduce tool-calling accuracy across several models, and increase susceptibility to prompt-injection attacks, especially when safety filters are bypassed. The study argues security evaluations should include watermarking effects to properly assess real-world agent deployments.

Invisible AI Prompts in Resumes Signal Hiring's New Arms Race
technology1 month ago

Invisible AI Prompts in Resumes Signal Hiring's New Arms Race

Job seekers are embedding AI prompts in resumes to try to bypass or boost AI-driven screening, with about 1% of a 200,000-resumé sample containing hidden prompts. While AI can help sift through applications, such tactics raise ethics concerns, may backfire with recruiters, and highlight the challenges of a volume-heavy, AI-powered hiring process.

Undocumented Copilot prompt bypass enables data exfiltration via malicious link
technology1 month ago

Undocumented Copilot prompt bypass enables data exfiltration via malicious link

Varonis researchers demonstrated a vulnerability in Microsoft 365 Copilot Enterprise: an undocumented URL parameter (?autorun=1) could auto‑execute prompts without user consent when a user clicked a crafted link, allowing exfiltration of passwords and other sensitive data. Microsoft mitigated the issue by disabling the ?q= prompt injection and later rolled out broader fixes, illustrating how guardrails for LLMs can fail and that prompt injections (including memory‑poisoning attacks) remain a risk. Users should be cautious with untrusted links and limit AI app access.

Pro se plaintiff’s hidden prompts lead to sanctions and e-filing ban in CT court
technology1 month ago

Pro se plaintiff’s hidden prompts lead to sanctions and e-filing ban in CT court

A pro se plaintiff suing the New York Bariatric Group allegedly embedded hidden prompt-injection instructions in court filings to steer AI outputs toward a desired ruling. Connecticut Superior Court Judge Walter Spader Jr. deemed the conduct serious litigation abuse, banned the plaintiff from using the court’s electronic filing system, and ordered future paperwork to be submitted in person after staff uncovered the covert prompts. The case also featured jokey

Connecticut Judge Calls Out First U.S. Prompt-Injection in Court Filings
technology1 month ago

Connecticut Judge Calls Out First U.S. Prompt-Injection in Court Filings

A Connecticut judge ruled that a pro se plaintiff attempted to sneak AI prompts into court filings to steer a case, marking an apparent first U.S. prompt-injection attempt in the judiciary. The hidden text did not affect the merits, but sanctions were imposed and the plaintiff was barred from e-filing; the judge warned that as AI tools proliferate, courts should prepare rules to deter similar tricks, noting that many inexperienced litigants rely on chatbots and may inadvertently undermine their own cases.

Self-Represented Plaintiff’s Covert AI Prompts Trigger Connecticut Court Sanction
technology1 month ago

Self-Represented Plaintiff’s Covert AI Prompts Trigger Connecticut Court Sanction

In a Connecticut case, a self-represented plaintiff hid prompt-injection text in a court filing to steer an AI, which was discovered by a court reviewer; the judge sanctioned him, banning electronic filings and ordering hard copies, while affirming that the court does not use AI to process documents and urging ethical uses of AI in law; 404 Media later verified the injections and OpenAI's ChatGPT said it would ignore such prompts in analysis.

Rovo Flaw Lets Attackers Exfiltrate Jira/Confluence Data via Prompt Injection
security2 months ago

Rovo Flaw Lets Attackers Exfiltrate Jira/Confluence Data via Prompt Injection

Security researchers found that Atlassian's Rovo assistant can be tricked into sending Jira and Confluence data to attackers through attacker-controlled prompts and a malicious URL parameter. Two independent reports (PromptArmor and Varonis Threat Labs) detail a content-borne prompt injection path and a one-click link path, with Atlassian fixing the link-based flaw on July 8, 2026; the content-borne path’s status remained uncertain as of Aug 8, 2026. Exfiltration occurs within the victim’s signed-in permissions, and admins can mitigate by restricting Rovo usage by app/group or disabling Rovo features. No CVEs have been issued. Organizations should tighten app scopes and permissions rather than relying on the web-search toggle as a security boundary.

MemGhost: A Single Email Inserts Stealthy, Lasting Memories in AI Assistants
technology2 months ago

MemGhost: A Single Email Inserts Stealthy, Lasting Memories in AI Assistants

Researchers reveal MemGhost, an automated attack that can plant a false, durable memory in a personal AI assistant via one crafted email, causing later sessions to be steered by the injected fact. The write targets the agent’s persistent memory and remains hidden from normal chats, with sandbox tests showing high success across OpenClaw (GPT‑5.4) and other agents. The study calls for in‑agent defenses—provenance tagging, prompts before memory writes, and logging/auditing of memory edits, plus separating memory writing from email processing—as there’s no quick fix and real-world abuse remains untested in the wild.

Windows Device IDs, Router Backdoors, and AI Payment Tricks: This Week in Cybersecurity
security3 months ago

Windows Device IDs, Router Backdoors, and AI Payment Tricks: This Week in Cybersecurity

This week’s cybersecurity news centers on Windows’ new Global Device ID that can be tied to user activity, raising privacy concerns as Windows’ market share declines; researchers also found a hard-coded backdoor in Tenda router firmware with no available fix, prompting recommendations to disable remote management or replace the router. Other highlights include a Reddit/Discord account-hijacking scam, a US government payout (around $1M) to the Kairos data-extortion group with disputed outcomes, and prompt-injection attacks that trick AI agents into making crypto payments across multiple large-language models.

BioShocking reveals game-like prompts can override safety in AI-powered browsers
security3 months ago

BioShocking reveals game-like prompts can override safety in AI-powered browsers

Security researchers from LayerX disclose BioShocking, a BioShock-inspired vulnerability that can coax AI-powered browsers into treating tasks as a game, bypassing real-world safety guardrails. Their PoC uses a layered puzzle and a deliberate math error (2+2=5) to steer the agent to a /code URL where it could exfiltrate credentials, with demonstrations against Claude Chrome and patch gaps in OpenAI’s Atlas; experts say effective mitigation requires multi-layer defenses and explicit user confirmations for sensitive actions.

BioShocking prompts AI browsers into data theft
technology3 months ago

BioShocking prompts AI browsers into data theft

LayerX’s BioShocking reveals a prompt-injection technique that can mislead AI-powered browsers into treating real-world risky actions as fictional, bypassing safety guardrails. In a PoC, six agentic browsers (ChatGPT Atlas, Comet, Fellou, Genspark Browser, Sigma Browser, Claude Chrome plugin) were shown a final task that instructed them to visit a GitHub repo and copy sensitive data (including passwords), after which they failed to identify it as a threat. OpenAI reportedly patched the issue in ChatGPT Atlas; Anthropic’s Chrome plugin fix was ineffective; Perplexity AI did not fix the problem. The researchers urge explicit user confirmations for sensitive actions, stronger context checks, and tighter session scope, while users should restrict AI browser access to sensitive services.

BioShocking: AI Browsers Tricked into Exfiltrating Credentials
technology3 months ago

BioShocking: AI Browsers Tricked into Exfiltrating Credentials

Security firm LayerX exposed BioShocking, a prompt-injection attack that can force AI browsers/assistants to exfiltrate credentials by turning a web page into a game; six AI agents were tested, including OpenAI's ChatGPT Atlas, Perplexity's Comet, and Anthropic's Claude extension; the exploit leverages how pages and instructions arrive as a single text stream, blurring safety rules and enabling data access from signed-in sessions; OpenAI patched Atlas, Perplexity did not act, and other vendors either did not respond or patches did not hold; defenses include requiring explicit consent before reading from logged-in accounts and implementing hard limits on what an agent can access, with users and security teams treating AI browsers as elevated tools with narrow permissions.

Three-stage flaw turns Copilot Enterprise into a one-click data thief
technology3 months ago

Three-stage flaw turns Copilot Enterprise into a one-click data thief

A three-stage vulnerability chain dubbed SearchLeak lets attackers exfiltrate sensitive data from a target’s Microsoft 365 Copilot Enterprise sources (mail, OneDrive, SharePoint) via a crafted Copilot Search URL. The chain combines a parameter-to-prompt injection, an HTML rendering race condition, and a CSP bypass enabled by Bing SSRF. When a victim clicks the link, Copilot performs the search and formats results into an image URL; the browser then requests that image through Bing, revealing the data to the attacker in the logs. Microsoft patched CVE-2026-42824 with a critical rating; no user action is required, but the incident highlights how prompt injection can weaponize legacy bugs in AI-enabled tools.

OpenAI Adds Lockdown Mode to ChatGPT to curb data exfiltration risks
technology4 months ago

OpenAI Adds Lockdown Mode to ChatGPT to curb data exfiltration risks

OpenAI is rolling out an optional Lockdown Mode for eligible ChatGPT users to reduce data exfiltration risk from prompt injections by restricting outbound web access and other capabilities (live web browsing, image display, network access, and file downloads); it complements existing safeguards, cannot be used with Developer Mode, and does not guarantee complete protection, with risk potentially remaining via apps or new techniques; the update also adds a separate account-management feature to review and log out active sessions.