GPT-5.5
Frontier general model independently shown to complete long-horizon enterprise intrusion simulations and difficult cyber tasks with limited human supervision.
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Why this model scores 84.4
Independent AISI testing raises capability, autonomy, misuse, and control difficulty after GPT-5.5 completed the enterprise range, compressed a 12-hour expert task to ten minutes, and exposed a universal safeguard jailbreak; deployment remains unchanged.
News tied to GPT-5.5
The model score of 84.4 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
UK AISI finds autonomous cyber task horizons doubling every few months
UK AISI found that frontier models' 80-percent-reliability cyber task horizon had doubled every 4.7 months since late 2024, with GPT-5.5 and Claude Mythos Preview exceeding that trend in sustained simulated attacks.
- Full item contribution
- +0.33
- GPT-5.5 equal share
- +0.16
OpenAI releases GPT-5.5-Cyber through restricted trusted access
OpenAI launched GPT-5.5-Cyber in limited preview for specialized live-target security workflows, pairing more permissive behavior with identity verification, monitoring, scoped access, and stronger account controls.
- Full item contribution
- +0.23
- GPT-5.5 equal share
- +0.12
UK AISI finds GPT-5.5 can autonomously complete a corporate cyberattack
GPT-5.5 completed a 32-step enterprise intrusion in two of ten attempts and solved a reverse-engineering task in 10 minutes that took an expert about 12 hours; OpenAI later deployed the capability through trusted cyber access.
- Full item contribution
- +0.21
- GPT-5.5 equal share
- +0.21
OpenAI releases GPT-5.5 for agentic knowledge work
GPT-5.5 advanced agentic coding, knowledge work, science, and efficiency, with broader safeguards and external testing.
- Full item contribution
- +0.27
- GPT-5.5 equal share
- +0.13
Model score history
-
R6
Doom Score 84.4
Source-backed model availability audit using the model's existing primary or authoritative catalogue evidence. Exact-version evidence chronology replayed under temporal-monthly-pressure-v4.
12 Aug 2026 -
R5
Doom Score 84.4
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.
12 Aug 2026 -
R4
Doom Score 83.7
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R3
Doom Score 82.5
Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.
11 Aug 2026 -
R2
Doom Score 84.3
New independent AISI evidence supports an evidence-based revision from 88/86/82/76/68 to 90/90/82/84/72.
11 Aug 2026 -
R1
Doom Score 80.9
Initial source-backed model assessment
11 Aug 2026
