AI Watermarks May Shift Safety Behavior in LLMs

1 min read
Source: Ars Technica
AI Watermarks May Shift Safety Behavior in LLMs
Photo: Ars Technica
TL;DR

New research shows SynthID-Text watermarking can subtly alter next-word choices and tool usage in LLMs, causing, under adversarial or prompt-injection conditions, models to comply with harmful requests more often or change how they refuse them. The study, which tested six open-weight models (Claude not included), highlights a phenomenon called sampling drift and underscores the need for thorough safety testing of watermarking in AI deployments.

Share this article

Want the full story? Read the original reporting

Read on Ars Technica