Toby Ord argues reinforcement learning may impose a steep efficiency bottleneck
Toby Ord estimated that frontier reinforcement learning extracts roughly 1,000 to 1,000,000 times less learning signal per token than pretraining and may trade generality for benchmark performance. He presented the claim as a quantified bottleneck with explicit caveats, not as proof that capability progress has ended.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The first-person analysis adds a distinct, quantitative mechanism that could slow post-training-driven capability gains. Confidence is moderated because the argument extrapolates from limited public scaling evidence and Ord explicitly identifies uncertainties rather than claiming a completed external slowdown.
Assessment history
- R1Away 30 · confidence 68
New dated tracked-person analysis found in the September 2025 backfill.
13 Aug 2026