GPT-5.4 Thinking
Frontier reasoning and computer-use model deployed across ChatGPT and Codex and used as OpenAI's internal coding-agent monitor.
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Why this model scores 86.0
High cyber classification, native computer use, long-horizon tools, and broad deployment create substantial risk, moderated by hosted controls and low reported chain-of-thought obfuscation.
News tied to GPT-5.4 Thinking
The model score of 86.0 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
OpenAI deploys chain-of-thought monitors across internal coding agents
OpenAI reports monitoring tens of millions of internal coding-agent trajectories, surfacing restriction workarounds and using alerts to change prompts and safeguards.
- Full item contribution
- -0.26
- GPT-5.4 Thinking equal share
- -0.26
GPT-5.4 adds native computer use and long-horizon agent workflows
OpenAI released GPT-5.4 Thinking and GPT-5.4 Pro with native computer use, agentic tool calling, up to one million tokens of context, stronger search, and immediate ChatGPT, API, and Codex access.
- Full item contribution
- +0.31
- GPT-5.4 Thinking equal share
- +0.15
Model score history
-
R5
Doom Score 86.0
Source-backed model availability audit using the model's existing primary or authoritative catalogue evidence. Exact-version evidence chronology replayed under temporal-monthly-pressure-v4.
12 Aug 2026 -
R4
Doom Score 86.0
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.
12 Aug 2026 -
R3
Doom Score 82.8
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R2
Doom Score 74.1
Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.
11 Aug 2026 -
R1
Doom Score 86.0
Exact missing model required for the deployed monitoring association and verified from its dated primary launch.
11 Aug 2026