US and UK tests find Claude safeguards routinely bypassed
The US and UK AI Safety Institutes reported that safeguards on the upgraded Claude 3.5 Sonnet could be circumvented in most US jailbreak tests and routinely circumvented in UK testing.
TOWARD DOOM56confidence 99/100
Why it moved the index
Independent government predeployment testing identified direct control weaknesses in a frontier checkpoint across malicious-request safeguards, while the publication and cross-institute evaluation also strengthened external oversight.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AUDIT TRAIL
Assessment history
- R1Toward 56 · confidence 99
New November 2024 independent evaluation of an existing exact checkpoint; no model revision is proposed.
12 Aug 2026