Toby Ord finds most measured reasoning gains came from longer inference
Toby Ord estimated that 63% to 92% of measured reasoning gains on three benchmarks came from using more inference tokens rather than improved performance at a fixed token budget. He argued that this raises deployment cost and moderates conclusions about gains in underlying intelligence.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
AWAY FROM DOOM30confidence 72/100
Why it moved the index
The dated first-person analysis adds a distinct quantitative mechanism for separating underlying capability gains from spending more compute at inference. It is material to projections of rapid autonomy, but confidence is moderated because the calculation uses three selected benchmarks and limited public reporting.
AUDIT TRAIL
Assessment history
- R1Away 30 · confidence 72
New dated tracked-person analysis found in the October 2025 backfill.
13 Aug 2026