xAI · Grok 4

Grok 4 Heavy

READER SUMMARY

Multi-agent Grok 4 tier using parallel test-time compute for difficult reasoning tasks.

DOOM SCORE77.5out of 100model risk profile, not the overall index
CURRENT ASSESSMENT · REVISION 1

Why this model scores 77.5

Parallel agents and scaled test-time compute raise capability, autonomy and control difficulty; expensive restricted access lowers deployment relative to Grok 4 but does not eliminate misuse potential.

Capability93
Autonomy86
Deployment31
Misuse potential84
Control difficulty72
0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

MODEL-ATTRIBUTED EVIDENCE

News tied to Grok 4 Heavy

The model score of 77.5 rates this model's risk profile. The overall Doom Index of 63.5 measures the recalibrated complete evidence corpus. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION+0.01

Each article's leave-one-out Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores are recalculated from their technical profile and related news, but never feed back into the overall index.

AUDIT TRAIL

Model score history

  1. R1
    Doom Score 77.5

    New separately named July 2025 multi-agent model profile verified by xAI.

    12 Aug 2026