Anthropic tests Constitutional Classifiers against universal jailbreaks
Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated over 300,000 interactions, ultimately finding one universal jailbreak and several narrower bypasses.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The method directly targets jailbreak pathways to chemical and biological harm and was exposed to large-scale adversarial testing. The score remains moderate because the live trial found a universal bypass and Anthropic had not yet deployed the method across production systems.
Assessment history
-
R1
Away 40 · confidence 94
New February 2025 safeguard result with practical live red-team evidence and no durable event collision.
12 Aug 2026
Share this page
-
DoomBench assesses “Anthropic tests Constitutional Classifiers against universal jailbreaks” as evidence moving away from doom, with magnitude 40 and confidence 94 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Anthropic tests Constitutional Classifiers against universal jailbreaks” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic tests Constitutional Classifiers against universal jailbreaks” as follows: Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated...
https://www.doombench.com/news/anthropic-tests-constitutional-classifiers-against-universal-jailbreaks-2025-02-03