Anthropic deploys nuclear-risk classifier on Claude traffic
Anthropic and the US National Nuclear Security Administration developed and deployed a classifier for nuclear-risk prompts, reporting 94.8 percent synthetic-query detection with no false positives in testing.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
A live classifier for nuclear-risk traffic creates a deployed safeguard against dangerous-domain access, with strong measured performance but developer-reported evaluation limits.
Assessment history
-
R1
Away 32 · confidence 80
New August 2025 deployed dangerous-domain safeguard; the moving Claude service does not establish an exact served checkpoint.
12 Aug 2026
Share this page
-
DoomBench assesses “Anthropic deploys nuclear-risk classifier on Claude traffic” as evidence moving away from doom, with magnitude 32 and confidence 80 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Anthropic deploys nuclear-risk classifier on Claude traffic” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic deploys nuclear-risk classifier on Claude traffic” as follows: Anthropic and the US National Nuclear Security Administration developed and deployed a classifier for nuclear-risk prompts, reporting 94.8...
https://www.doombench.com/news/anthropic-deploys-nuclear-risk-classifier-on-claude-traffic-2025-08-21