Anthropic adopts Responsible Scaling Policy with capability thresholds
Anthropic adopted a board-approved policy tying stronger security and deployment safeguards to capability thresholds, including a commitment not to deploy ASL-3 systems without adequate measures.
AWAY FROM DOOM51confidence 96/100
Why it moved the index
The policy created explicit capability-triggered deployment and security gates for catastrophic misuse risks, establishing a concrete frontier scaling control despite being self-imposed.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AUDIT TRAIL
Assessment history
- R1Away 51 · confidence 96
New September 2023 capability-triggered frontier safety policy.
12 Aug 2026