Safety and alignment

Anthropic strengthens catastrophic-risk scaling policy

Anthropic adopted Responsible Scaling Policy version 2.0 with thresholds for autonomous AI research and CBRN assistance, escalating security and deployment safeguards as capabilities advance.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM48confidence 99/100

Why it moved the index

A board-governed policy linked explicit catastrophic capability thresholds to stronger controls and a commitment not to train or deploy without adequate safeguards, while remaining an internally enforced framework.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 48 · confidence 99

    New October 2024 frontier catastrophic-risk control framework with no durable collision.

    12 Aug 2026