Safety and alignment

OpenAI and Anthropic publish joint model safety evaluations

OpenAI and Anthropic jointly evaluated each other's frontier models across alignment, misuse and capability tests, publishing strengths, failures and limits of extrapolating the results to real-world behavior.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM28confidence 82/100

Why it moved the index

Cross-laboratory evaluation exposed concrete safety strengths and failures across exact frontier checkpoints, improving independent scrutiny while remaining evaluation evidence rather than a deployed control.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 28 · confidence 82

    New August 2025 cross-lab safety evaluation with exact model relationships and no durable collision.

    12 Aug 2026