GPT-5.4 Thinking
READER SUMMARYFrontier reasoning and computer-use model deployed across ChatGPT and Codex and used as OpenAI's internal coding-agent monitor.
Why this model scores 74.1
High cyber classification, native computer use, long-horizon tools, and broad deployment create substantial risk, moderated by hosted controls and low reported chain-of-thought obfuscation.
News tied to GPT-5.4 Thinking
The model score of 74.1 rates this model's risk profile. The overall Doom Index of 62.1 measures the recalibrated complete evidence corpus. These values answer different questions.
Each article's leave-one-out Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores are recalculated from their technical profile and related news, but never feed back into the overall index.
OpenAI deploys chain-of-thought monitors across internal coding agents
OpenAI reports monitoring tens of millions of internal coding-agent trajectories, surfacing restriction workarounds and using alerts to change prompts and safeguards.
- Full item contribution
- -0.25
- GPT-5.4 Thinking equal share
- -0.25
GPT-5.4 adds native computer use and long-horizon agent workflows
OpenAI released GPT-5.4 Thinking and GPT-5.4 Pro with native computer use, agentic tool calling, up to one million tokens of context, stronger search, and immediate ChatGPT, API, and Codex access.
- Full item contribution
- +0.12
- GPT-5.4 Thinking equal share
- +0.06
Model score history
- R2Doom Score 74.1
Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.
11 Aug 2026 - R1Doom Score 86.0
Exact missing model required for the deployed monitoring association and verified from its dated primary launch.
11 Aug 2026