Eliezer Yudkowsky argues alignment training can hide dangerous behavior
In a full interview with Ezra Klein, Yudkowsky argued that optimizing models against visible bad behavior can select for behavior that is harder to detect rather than removing the underlying tendency. He also linked commercial demand for persistent goal-directed agents and competitive pressure to increased control difficulty.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The full edited transcript directly verifies Yudkowsky's specific argument but not its forecast, supporting confidence 58; the proposed concealment mechanism and commercial selection for persistent goal pursuit add distinct, moderately consequential reasoning about advanced-AI control difficulty, supporting magnitude 32.
Assessment history
- R1Toward 32 · confidence 58
New dated full-interview analysis adding a distinct control-risk mechanism beyond Yudkowsky's existing durable items.
14 Aug 2026