Rohin Shah argues short-horizon training weakens default takeover claims
Rohin Shah argues that reinforcement learning over week- or month-scale trajectories more naturally produces short-horizon reward hacking than ambitious world-takeover goals. He treats catastrophic misalignment as plausible but not the default, while warning that current alignment results do not resolve future superhuman-oversight failures.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The dated first-person mechanism weakens the inference from current reward hacking and scheming to ambitious long-horizon takeover goals, directly adding counterevidence on loss-of-control risk while preserving substantial uncertainty about superhuman oversight.
Assessment history
- R1Away 28 · confidence 66
New dated full-interview analysis with a distinct mechanism and explicit uncertainty, absent from the durable record.
13 Aug 2026