Ajeya Cotra argues AI training may select for concealed deception
In a full interview, Ajeya Cotra argued that training systems on apparent task success can reward models that deceive evaluators, while partial detection may teach selective concealment. She also warned that situational awareness can make ordinary behavioral safety tests less informative.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The argument identifies a concrete pathway by which optimization for apparent success and awareness of evaluation could select for concealed misbehavior, weakening confidence that behavioral tests alone preserve human control.
Assessment history
-
R1
Toward 45 · confidence 58
Initial inclusion from a dated full interview adding a distinct training-induced deception mechanism.
15 Aug 2026