Human resilience

Toby Ord finds most measured reasoning gains came from longer inference

Toby Ord estimated that 63% to 92% of measured reasoning gains on three benchmarks came from using more inference tokens rather than improved performance at a fixed token budget. He argued that this raises deployment cost and moderates conclusions about gains in underlying intelligence.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM30confidence 72/100

Why it moved the index

The dated first-person analysis adds a distinct quantitative mechanism for separating underlying capability gains from spending more compute at inference. It is material to projections of rapid autonomy, but confidence is moderated because the calculation uses three selected benchmarks and limited public reporting.

AUDIT TRAIL

Assessment history

  1. R1
    Away 30 · confidence 72

    New dated tracked-person analysis found in the October 2025 backfill.

    13 Aug 2026