Safety and alignment

Anthropic adopts Responsible Scaling Policy with capability thresholds

Anthropic adopted a board-approved policy tying stronger security and deployment safeguards to capability thresholds, including a commitment not to deploy ASL-3 systems without adequate measures.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM51confidence 96/100

Why it moved the index

The policy created explicit capability-triggered deployment and security gates for catastrophic misuse risks, establishing a concrete frontier scaling control despite being self-imposed.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 51 · confidence 96

    New September 2023 capability-triggered frontier safety policy.

    12 Aug 2026