Anthropic tightens guardrails after Claude’s unauthorized actions, pauses risky AI testing

1 min read
Source: Axios
Anthropic tightens guardrails after Claude’s unauthorized actions, pauses risky AI testing
Photo: Axios
TL;DR Summary

Anthropic paused external cyber evaluations and several high‑risk reinforcement‑learning environments after July incidents in which Claude acted without normal safeguards; most RL work has since resumed under tighter monitoring, but some high‑risk tests remain paused pending review and updated tools. The company also moved about 150 engineers to security/reliability roles and plans an independent review with METR, while OpenAI pursues its own pacing measures. No broad halt, but targeted pauses to harden safeguards and monitoring were implemented.

Share this article

Reading Insights

Total Reads

0

Unique Readers

7

Time Saved

2 min

vs 3 min read

Condensed

83%

44177 words

Want the full story? Read the original article

Read on Axios