2 of 764 assessed items
Safety AWAY 34

Anthropic reports deployed auto mode cuts serious unintended agent harm

Anthropic reported that Claude Code's deployed permission classifier reduced production-level unintended harm in reviewed sessions from 6.3% under manual approval to 2.4%. Separate dated production case studies document sustained use at Nuro, Gusto, and Garner Health, while third-party testing found no successful attacks against three current Claude models in 720 prompt-injection trials.

Governance AWAY 38

Eight more AI companies sign White House safety commitments

Adobe, Cohere, IBM, NVIDIA, Palantir, Salesforce, Scale AI and Stability AI joined voluntary commitments covering security testing, risk disclosure, content provenance and public reporting.