Safety and alignment

Anthropic tests Constitutional Classifiers against universal jailbreaks

Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated over 300,000 interactions, ultimately finding one universal jailbreak and several narrower bypasses.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM40confidence 94/100

Why it moved the index

The method directly targets jailbreak pathways to chemical and biological harm and was exposed to large-scale adversarial testing. The score remains moderate because the live trial found a universal bypass and Anthropic had not yet deployed the method across production systems.

AUDIT TRAIL

Assessment history

  1. R1
    Away 40 · confidence 94

    New February 2025 safeguard result with practical live red-team evidence and no durable event collision.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic tests Constitutional Classifiers against universal jailbreaks.
  1. DoomBench assesses “Anthropic tests Constitutional Classifiers against universal jailbreaks” as evidence moving away from doom, with magnitude 40 and confidence 94 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic tests Constitutional Classifiers against universal jailbreaks” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic tests Constitutional Classifiers against universal jailbreaks” as follows: Anthropic published and live-tested a classifier-based safeguard for dangerous-content jailbreaks; 339 participants generated...

    https://www.doombench.com/news/anthropic-tests-constitutional-classifiers-against-universal-jailbreaks-2025-02-03