← Model benchmark
OpenAI · GPT

GPT-4

A multimodal frontier model that reached human-level performance on several professional and academic benchmarks while retaining material reliability limits.

DOOM SCORE51.4out of 100
CURRENT ASSESSMENT · REVISION 1

Why this model scores 51.4

GPT-4 was a major general-capability step with broad API deployment, but it lacked the sustained agentic autonomy and dangerous-domain performance of later frontier systems.

Capability64
Autonomy32
Deployment68
Misuse potential48
Control difficulty42
AUDIT TRAIL

Model score history

  1. R1
    Doom Score 51.4

    Initial source-backed model assessment

    11 Aug 2026