Safety and alignment

Anthropic finds long contexts enable many-shot jailbreaking

Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the weakness to peers and deployed prompt-classification mitigations that reduced one measured attack rate from 61 percent to 2 percent.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM54confidence 96/100

Why it moved the index

The research exposed a simple cross-provider method for bypassing long-context safeguards, and later Anthropic system-card evidence confirmed continued susceptibility in a released frontier model despite external safety layers.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 54 · confidence 96

    New April 2024 safeguard-weakness result paired with verified deployed-model impact and mitigation evidence.

    12 Aug 2026