Safety and alignment

Anthropic maps and steers safety-relevant features in Claude 3 Sonnet

Anthropic scaled sparse feature extraction to the publicly deployed Claude 3 Sonnet, identified internal representations tied to unsafe code, bias, deception and manipulation, and demonstrated that activating or suppressing features could causally change model behavior.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM43confidence 86/100

Why it moved the index

The result moved mechanistic interpretation from toy systems to a released frontier model and produced direct causal steering of safety-relevant behaviors. Dated independent reporting verified the released-model intervention while also documenting that reliable production safety control remained unproven and expensive.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 43 · confidence 86

    New May 2024 primary safety result paired with dated authoritative evidence of practical behavior steering in a released model.

    12 Aug 2026