Safety and alignment

OpenAI releases open-weight safety reasoners backed by deployed classifier practice

OpenAI released 120B and 20B open-weight safeguard models that interpret custom policies, with the underlying safety-reasoning approach already used in production systems.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 88/100

Why it moved the index

The release gives developers inspectable and adaptable policy-enforcement models based on a safety-reasoning approach already used in production systems.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 88

    New source-verified safeguard models and practical deployment evidence absent from durable context.

    11 Aug 2026