2 of 825 assessed items
Resilience AWAY 28

Rohin Shah argues short-horizon training weakens default takeover claims

Rohin Shah argues that reinforcement learning over week- or month-scale trajectories more naturally produces short-horizon reward hacking than ambitious world-takeover goals. He treats catastrophic misalignment as plausible but not the default, while warning that current alignment results do not resolve future superhuman-oversight failures.

Governance TOWARD 20

Rohin Shah warns a short AI pause could reduce safety time

Rohin Shah supports enforceable global slowing, but identifies a countervailing mechanism: hardware, algorithms, and investment could keep advancing during a training pause while safety researchers lose access to intermediate models. A poorly designed pause could therefore reduce capabilities-adjusted safety time without proportionally delaying dangerous systems.