OpenAI · GPT
GPT-4
A multimodal frontier model that reached human-level performance on several professional and academic benchmarks while retaining material reliability limits.
DOOM SCORE51.4out of 100
Why this model scores 51.4
GPT-4 was a major general-capability step with broad API deployment, but it lacked the sustained agentic autonomy and dangerous-domain performance of later frontier systems.
Capability64
Autonomy32
Deployment68
Misuse potential48
Control difficulty42
AUDIT TRAIL
Model score history
- R1Doom Score 51.4
Initial source-backed model assessment
11 Aug 2026