Anthropic finds long contexts enable many-shot jailbreaking
Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the weakness to peers and deployed prompt-classification mitigations that reduced one measured attack rate from 61 percent to 2 percent.
TOWARD DOOM54confidence 96/100
Why it moved the index
The research exposed a simple cross-provider method for bypassing long-context safeguards, and later Anthropic system-card evidence confirmed continued susceptibility in a released frontier model despite external safety layers.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AUDIT TRAIL
Assessment history
- R1Toward 54 · confidence 96
New April 2024 safeguard-weakness result paired with verified deployed-model impact and mitigation evidence.
12 Aug 2026