Safety and alignment

Anthropic deploys conversation-ending safeguards in Claude

Anthropic enabled Claude Opus 4 and Opus 4.1 to end a narrow subset of persistently harmful conversations after other response strategies fail.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM22confidence 76/100

Why it moved the index

A deployed model-initiated stop mechanism adds a concrete control against sustained harmful interactions, though evidence is limited to the developer's narrow self-report.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 22 · confidence 76

    New August 2025 deployed safeguard affecting two exact Claude checkpoints; no durable collision.

    12 Aug 2026