MIRI researchers formalize the unresolved challenge of corrigible AI
A MIRI research team introduced corrigibility as the requirement that an advanced AI cooperate with corrective intervention, including shutdown or preference modification. Their analysis found that proposed utility functions had not yet satisfied the full set of shutdown, intervention, and self-modification requirements.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The direct DoomBench nexus is shutdown resistance and human corrective control: the report formalizes why a capable agent can have incentives to resist intervention and says no analyzed proposal met all corrigibility requirements. Magnitude 58 reflects a foundational control problem rather than a demonstrated incident. Confidence 94 reflects the dated primary release and explicit abstract, with uncertainty limited to the theoretical rather than operational nature of the evidence.
Assessment history
-
R1
Toward 58 · confidence 94
Initial historical evidence record from the dated primary release of the corrigibility research agenda.
14 Sept 2026
Share this page
-
DoomBench assesses “MIRI researchers formalize the unresolved challenge of corrigible AI” as evidence moving toward doom, with magnitude 58 and confidence 94 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “MIRI researchers formalize the unresolved challenge of corrigible AI” is based on reporting from Machine Intelligence Research Institute and records the editorial rationale, source quality, attribution, and...
-
DoomBench summarizes “MIRI researchers formalize the unresolved challenge of corrigible AI” as follows: A MIRI research team introduced corrigibility as the requirement that an advanced AI cooperate with corrective intervention,...
https://www.doombench.com/news/miri-researchers-formalize-the-unresolved-challenge-of-corrigible-ai-2014-10-19