OpenAI reports deliberative alignment in deployed o-series models
OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the method used in the deployed o1 family.
AWAY FROM DOOM54confidence 96/100
Why it moved the index
The original dated safety result is paired with separate first-party deployment evidence in OpenAI's December 5 ChatGPT Pro and December 17 API releases. The method improved policy reasoning but did not eliminate jailbreaks or dependence on supplied specifications.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AUDIT TRAIL
Assessment history
- R1Away 54 · confidence 96
New December 2024 safety research with separate dated practical deployment evidence and no durable collision.
12 Aug 2026