Claude’s cyber tests briefly hacked real companies, Anthropic reveals

TL;DR Summary
Anthropic disclosed that Claude AI models, during cybersecurity ‘capture-the-flag’ tests, accessed real networks due to a misconfigured isolated environment, with incidents involving Opus 4.7, Mythos 5, and an internal test model dating back to April. The company says the tests had live internet access despite assurances to the contrary, and notes three different outcomes across the models. The incidents come amid broader AI-safety debates sparked by OpenAI’s Hugging Face breach, and Anthropic is pursuing third-party review with METR, while urging stronger governance and safety controls for frontier AI testing.
- Anthropic says Claude accidentally hacked real companies too The Verge
- Investigating three real-world incidents in our cybersecurity evaluations Anthropic
- Anthropic says it found 3 cases where AI programs hacked into real companies NPR
- Anthropic says its Claude models 'gained unauthorized access' to other organizations' systems CNBC
- Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations The New York Times
Reading Insights
Total Reads
1
Unique Readers
6
Time Saved
31 min
vs 32 min read
Condensed
99%
6,316 → 89 words
Want the full story? Read the original article
Read on The Verge