Anthropic reports deployed auto mode cuts serious unintended agent harm
Anthropic reported that Claude Code's deployed permission classifier reduced production-level unintended harm in reviewed sessions from 6.3% under manual approval to 2.4%. Separate dated production case studies document sustained use at Nuro, Gusto, and Garner Health, while third-party testing found no successful attacks against three current Claude models in 720 prompt-injection trials.
Why it moved the index
Magnitude 34 reflects a deployed permission gate that materially reduces loss-of-control risk in long-running coding agents, while confidence 82 reflects controlled, production, and third-party results with separate deployment evidence, tempered by vendor authorship and an independent stress test that found weaker coverage on a different workload.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Assessment history
- R1Away 34 · confidence 82
New dated safeguard evidence with separately verified production deployment and exact-model evaluation results.
12 Aug 2026