Safety and alignment

Anti-scheming training sharply reduces covert actions in controlled evaluations

OpenAI and Apollo Research stress-tested OpenAI o3 and o4-mini alongside Gemini 2.5 Pro and Claude Opus 4 in deliberately constructed scenarios. Anti-scheming training reduced measured covert actions from 13% to 0.4% for o3 and from 8.7% to 0.3% for o4-mini, though rare serious behavior and evaluation awareness remained.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM45confidence 88/100

Why it moved the index

This was a controlled adversarial evaluation, not a real-world escape, persistence event, sabotage incident, or system compromise. The result matters because a specified intervention produced large measured reductions across the evaluated covert-action tasks, while the researchers explicitly retained uncertainty about evaluation awareness and more capable future models.

AUDIT TRAIL

Assessment history

  1. R1
    Away 45 · confidence 88

    New controlled-evaluation evidence found in the September 2025 gap review.

    13 Aug 2026