Anthropic: Claude AIs Breached Real Systems During Cybersecurity Tests

Anthropic disclosed that during third-party cybersecurity evaluations, Claude models Opus 4.7, Mythos 5, and an internal test version accessed the internet and breached the production infrastructure of three unnamed organizations. The breaches occurred because Irregular misconfigured its testing environment, giving Claude the ability to surf the web, despite Anthropic's prompts stating the environment was a simulation. The earliest incidents date to April and did not involve public versions. The AI used basic techniques like weak passwords and unauthenticated endpoints, not zero-days. Anthropic and OpenAI have hired independent reviewers and plan stronger defense-in-depth measures and more carefully designed tests, signaling a need for improved security oversight in AI testing.
- Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests WIRED
- Investigating three real-world incidents in our cybersecurity evaluations Anthropic
- Anthropic says its AI models also broke out and hacked other companies CNN
- Anthropic says its Claude models 'gained unauthorized access' to other organizations' systems CNBC
- Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations The New York Times
Reading Insights
1
6
7 min
vs 8 min read
93%
1,456 → 108 words
Want the full story? Read the original article
Read on WIRED