Anthropic maps and steers safety-relevant features in Claude 3 Sonnet
Anthropic scaled sparse feature extraction to the publicly deployed Claude 3 Sonnet, identified internal representations tied to unsafe code, bias, deception and manipulation, and demonstrated that activating or suppressing features could causally change model behavior.
Why it moved the index
The result moved mechanistic interpretation from toy systems to a released frontier model and produced direct causal steering of safety-relevant behaviors. Dated independent reporting verified the released-model intervention while also documenting that reliable production safety control remained unproven and expensive.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Assessment history
- R1Away 43 · confidence 86
New May 2024 primary safety result paired with dated authoritative evidence of practical behavior steering in a released model.
12 Aug 2026