Safety and alignment

Anthropic adopts Responsible Scaling Policy with capability thresholds

Anthropic adopted a board-approved policy tying stronger security and deployment safeguards to capability thresholds, including a commitment not to deploy ASL-3 systems without adequate measures.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM51confidence 96/100

Why it moved the index

The policy created explicit capability-triggered deployment and security gates for catastrophic misuse risks, establishing a concrete frontier scaling control despite being self-imposed.

AUDIT TRAIL

Assessment history

  1. R1
    Away 51 · confidence 96

    New September 2023 capability-triggered frontier safety policy.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic adopts Responsible Scaling Policy with capability thresholds.
  1. DoomBench assesses “Anthropic adopts Responsible Scaling Policy with capability thresholds” as evidence moving away from doom, with magnitude 51 and confidence 96 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic adopts Responsible Scaling Policy with capability thresholds” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic adopts Responsible Scaling Policy with capability thresholds” as follows: Anthropic adopted a board-approved policy tying stronger security and deployment safeguards to capability thresholds, including a...

    https://www.doombench.com/news/anthropic-adopts-responsible-scaling-policy-with-capability-thresholds-2023-09-19