Safety and alignment

Anthropic deploys nuclear-risk classifier on Claude traffic

Anthropic and the US National Nuclear Security Administration developed and deployed a classifier for nuclear-risk prompts, reporting 94.8 percent synthetic-query detection with no false positives in testing.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM32confidence 80/100

Why it moved the index

A live classifier for nuclear-risk traffic creates a deployed safeguard against dangerous-domain access, with strong measured performance but developer-reported evaluation limits.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 32 · confidence 80

    New August 2025 deployed dangerous-domain safeguard; the moving Claude service does not establish an exact served checkpoint.

    12 Aug 2026