Safety and alignment

Anthropic finds Claude's expressed values vary across models and languages

Anthropic analyzed 309,815 production conversations and found structured differences in Claude's expressed values across model versions and the twenty most common languages. The variations were not deliberately chosen, and the team proposed using the measurement method for pre-release evaluation and post-deployment monitoring while noting that user impacts remain unmeasured.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM25confidence 83/100

Why it moved the index

The deployment-scale analysis provides direct evidence that model behavior varies across versions and languages in ways not intentionally selected, which complicates consistent control and monitoring. The effect sizes are structured but modest, only 15% of value variation is captured, and downstream user harm was not measured.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 25 · confidence 83

    Adds deployment-scale evidence of unintended value variation across exact Claude versions and twenty languages.

    15 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic finds Claude's expressed values vary across models and languages.
  1. DoomBench assesses “Anthropic finds Claude's expressed values vary across models and languages” as evidence moving toward doom, with magnitude 25 and confidence 83 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic finds Claude's expressed values vary across models and languages” is based on reporting from Anthropic Research and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic finds Claude's expressed values vary across models and languages” as follows: Anthropic analyzed 309,815 production conversations and found structured differences in Claude's expressed values across model...

    https://www.doombench.com/news/anthropic-finds-claude-s-expressed-values-vary-across-models-and-languages-2026-07-13