Safety and alignment

OpenAI reports deliberative alignment in deployed o-series models

OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the method used in the deployed o1 family.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM54confidence 96/100

Why it moved the index

The original dated safety result is paired with separate first-party deployment evidence in OpenAI's December 5 ChatGPT Pro and December 17 API releases. The method improved policy reasoning but did not eliminate jailbreaks or dependence on supplied specifications.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 54 · confidence 96

    New December 2024 safety research with separate dated practical deployment evidence and no durable collision.

    12 Aug 2026