Human resilience

Toby Ord argues reinforcement learning may impose a steep efficiency bottleneck

Toby Ord estimated that frontier reinforcement learning extracts roughly 1,000 to 1,000,000 times less learning signal per token than pretraining and may trade generality for benchmark performance. He presented the claim as a quantified bottleneck with explicit caveats, not as proof that capability progress has ended.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM30confidence 68/100

Why it moved the index

The first-person analysis adds a distinct, quantitative mechanism that could slow post-training-driven capability gains. Confidence is moderated because the argument extrapolates from limited public scaling evidence and Ord explicitly identifies uncertainties rather than claiming a completed external slowdown.

AUDIT TRAIL

Assessment history

  1. R1
    Away 30 · confidence 68

    New dated tracked-person analysis found in the September 2025 backfill.

    13 Aug 2026