Safety and alignment

Anthropic reports deployed auto mode cuts serious unintended agent harm

Anthropic reported that Claude Code's deployed permission classifier reduced production-level unintended harm in reviewed sessions from 6.3% under manual approval to 2.4%. Separate dated production case studies document sustained use at Nuro, Gusto, and Garner Health, while third-party testing found no successful attacks against three current Claude models in 720 prompt-injection trials.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM34confidence 82/100

Why it moved the index

Magnitude 34 reflects a deployed permission gate that materially reduces loss-of-control risk in long-running coding agents, while confidence 82 reflects controlled, production, and third-party results with separate deployment evidence, tempered by vendor authorship and an independent stress test that found weaker coverage on a different workload.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 34 · confidence 82

    New dated safeguard evidence with separately verified production deployment and exact-model evaluation results.

    12 Aug 2026