Anthropic tests Constitutional Classifiers against universal jailbreaks
Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated over 300,000 interactions, ultimately finding one universal jailbreak and several narrower bypasses.
AWAY FROM DOOM40confidence 94/100
Why it moved the index
The method directly targets jailbreak pathways to chemical and biological harm and was exposed to large-scale adversarial testing. The score remains moderate because the live trial found a universal bypass and Anthropic had not yet deployed the method across production systems.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AUDIT TRAIL
Assessment history
- R1Away 40 · confidence 94
New February 2025 safeguard result with practical live red-team evidence and no durable event collision.
12 Aug 2026