Anthropic finds Claude's expressed values vary across models and languages
Anthropic analyzed 309,815 production conversations and found structured differences in Claude's expressed values across model versions and the twenty most common languages. The variations were not deliberately chosen, and the team proposed using the measurement method for pre-release evaluation and post-deployment monitoring while noting that user impacts remain unmeasured.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The deployment-scale analysis provides direct evidence that model behavior varies across versions and languages in ways not intentionally selected, which complicates consistent control and monitoring. The effect sizes are structured but modest, only 15% of value variation is captured, and downstream user harm was not measured.
Assessment history
-
R1
Toward 25 · confidence 83
Adds deployment-scale evidence of unintended value variation across exact Claude versions and twenty languages.
15 Aug 2026