Safety and alignment

Anthropic tests Constitutional Classifiers against universal jailbreaks

Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated over 300,000 interactions, ultimately finding one universal jailbreak and several narrower bypasses.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM40confidence 94/100

Why it moved the index

The method directly targets jailbreak pathways to chemical and biological harm and was exposed to large-scale adversarial testing. The score remains moderate because the live trial found a universal bypass and Anthropic had not yet deployed the method across production systems.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 40 · confidence 94

    New February 2025 safeguard result with practical live red-team evidence and no durable event collision.

    12 Aug 2026