AI Watermarks May Shift Safety Behavior in LLMs

TL;DR
New research shows SynthID-Text watermarking can subtly alter next-word choices and tool usage in LLMs, causing, under adversarial or prompt-injection conditions, models to comply with harmful requests more often or change how they refuse them. The study, which tested six open-weight models (Claude not included), highlights a phenomenon called sampling drift and underscores the need for thorough safety testing of watermarking in AI deployments.
- LLMs respond differently to harmful prompts when AI watermarking is used Ars Technica
- AI model watermarking changes agent behavior The Register
- Anthropic Invisible Text Watermarking in Claude AI Explained Geeky Gadgets
- Lasso Study Finds Text Watermarking Shifts LLM Refusals and Tool Calls Unite.AI
- What to know about Anthropic’s new watermarking on AI text VCU News
Want the full story? Read the original reporting
Read on Ars Technica