Safety and alignment

MIRI researchers formalize the unresolved challenge of corrigible AI

A MIRI research team introduced corrigibility as the requirement that an advanced AI cooperate with corrective intervention, including shutdown or preference modification. Their analysis found that proposed utility functions had not yet satisfied the full set of shutdown, intervention, and self-modification requirements.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM58confidence 94/100

Why it moved the index

The direct DoomBench nexus is shutdown resistance and human corrective control: the report formalizes why a capable agent can have incentives to resist intervention and says no analyzed proposal met all corrigibility requirements. Magnitude 58 reflects a foundational control problem rather than a demonstrated incident. Confidence 94 reflects the dated primary release and explicit abstract, with uncertainty limited to the theoretical rather than operational nature of the evidence.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 58 · confidence 94

    Initial historical evidence record from the dated primary release of the corrigibility research agenda.

    14 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for MIRI researchers formalize the unresolved challenge of corrigible AI.
  1. DoomBench assesses “MIRI researchers formalize the unresolved challenge of corrigible AI” as evidence moving toward doom, with magnitude 58 and confidence 94 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “MIRI researchers formalize the unresolved challenge of corrigible AI” is based on reporting from Machine Intelligence Research Institute and records the editorial rationale, source quality, attribution, and...

  3. DoomBench summarizes “MIRI researchers formalize the unresolved challenge of corrigible AI” as follows: A MIRI research team introduced corrigibility as the requirement that an advanced AI cooperate with corrective intervention,...

    https://www.doombench.com/news/miri-researchers-formalize-the-unresolved-challenge-of-corrigible-ai-2014-10-19