Victoria Krakovna maps AI specification failures to four Goodhart effects
Victoria Krakovna and Ramana Kumar map objective-specification failures to regressional, extremal, causal, and adversarial Goodhart effects. Their framework distinguishes gaps between ideal, model, design, implementation, and revealed objectives, linking them to specification gaming, side effects, reward tampering, robustness failures, and deceptive alignment.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The analysis adds a distinct mechanism-level taxonomy showing how optimizing imperfect proxies can produce specification gaming, side effects, tampering, robustness failures, and deceptive alignment through different objective gaps. It clarifies control difficulty but does not claim that these failure modes occurred in deployed systems.
Assessment history
-
R1
Toward 22 · confidence 78
Adds a previously absent, exactly timestamped taxonomy connecting objective gaps to multiple AI control-failure mechanisms.
22 Sept 2026
Share this page
-
DoomBench assesses “Victoria Krakovna maps AI specification failures to four Goodhart effects” as evidence moving toward doom, with magnitude 22 and confidence 78 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Victoria Krakovna maps AI specification failures to four Goodhart effects” is based on reporting from Victoria Krakovna and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Victoria Krakovna maps AI specification failures to four Goodhart effects” as follows: Victoria Krakovna and Ramana Kumar map objective-specification failures to regressional, extremal, causal, and adversarial...
https://www.doombench.com/news/victoria-krakovna-maps-ai-specification-failures-to-four-goodhart-effects-2019-08-19