Safety and alignment

Anthropic deploys nuclear-risk classifier on Claude traffic

Anthropic and the US National Nuclear Security Administration developed and deployed a classifier for nuclear-risk prompts, reporting 94.8 percent synthetic-query detection with no false positives in testing.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM32confidence 80/100

Why it moved the index

A live classifier for nuclear-risk traffic creates a deployed safeguard against dangerous-domain access, with strong measured performance but developer-reported evaluation limits.

AUDIT TRAIL

Assessment history

  1. R1
    Away 32 · confidence 80

    New August 2025 deployed dangerous-domain safeguard; the moving Claude service does not establish an exact served checkpoint.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic deploys nuclear-risk classifier on Claude traffic.
  1. DoomBench assesses “Anthropic deploys nuclear-risk classifier on Claude traffic” as evidence moving away from doom, with magnitude 32 and confidence 80 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic deploys nuclear-risk classifier on Claude traffic” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic deploys nuclear-risk classifier on Claude traffic” as follows: Anthropic and the US National Nuclear Security Administration developed and deployed a classifier for nuclear-risk prompts, reporting 94.8...

    https://www.doombench.com/news/anthropic-deploys-nuclear-risk-classifier-on-claude-traffic-2025-08-21