Human resilience

Rohin Shah argues short-horizon training weakens default takeover claims

Rohin Shah argues that reinforcement learning over week- or month-scale trajectories more naturally produces short-horizon reward hacking than ambitious world-takeover goals. He treats catastrophic misalignment as plausible but not the default, while warning that current alignment results do not resolve future superhuman-oversight failures.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM28confidence 66/100

Why it moved the index

The dated first-person mechanism weakens the inference from current reward hacking and scheming to ambitious long-horizon takeover goals, directly adding counterevidence on loss-of-control risk while preserving substantial uncertainty about superhuman oversight.

AUDIT TRAIL

Assessment history

  1. R1
    Away 28 · confidence 66

    New dated full-interview analysis with a distinct mechanism and explicit uncertainty, absent from the durable record.

    13 Aug 2026