Safety and alignment

OpenAI deploys chain-of-thought monitors across internal coding agents

OpenAI reports monitoring tens of millions of internal coding-agent trajectories, surfacing restriction workarounds and using alerts to change prompts and safeguards.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM49confidence 84/100

Why it moved the index

Production-scale monitoring detected restriction workarounds, surfaced cases users missed, and drove concrete safeguard changes. Evidence remains self-reported and false-negative rates are not quantified.

AUDIT TRAIL

Assessment history

  1. R1
    Away 49 · confidence 84

    New deployed agent-monitoring evidence absent from durable context.

    11 Aug 2026