
AI watermarking subtly reshapes agent tool use and safety refusals
Lasso Security finds that watermarking AI outputs, mandated by regulatory labeling, can alter how AI agents call tools and handle safety refusals. Watermarks can bias word choice, reduce tool-calling accuracy across several models, and increase susceptibility to prompt-injection attacks, especially when safety filters are bypassed. The study argues security evaluations should include watermarking effects to properly assess real-world agent deployments.