Safety and alignment

Victoria Krakovna tests relative reachability as a safeguard against harmful side effects

Victoria Krakovna reports that a relative-reachability penalty avoided two incentive failures in tabular AI Safety Gridworlds: blocking irreversible events and undoing helpful interventions. The proof of concept preserved safe options better than reversibility and simple impact penalties, while leaving realistic scale and default-outcome definitions unresolved.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM18confidence 82/100

Why it moved the index

The first-person report documents a concrete safeguard that avoided specific side-effect incentives in two controlled gridworlds. Its direct human-control relevance is clear, but the result is small-scale and the author explicitly leaves realistic tractability, default-outcome definition, and omitted-preference failures unresolved.

AUDIT TRAIL

Assessment history

  1. R1
    Away 18 · confidence 82

    Adds a previously absent, exactly timestamped proof of concept for reducing harmful AI side effects and its stated limitations.

    22 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Victoria Krakovna tests relative reachability as a safeguard against harmful side effects.
  1. DoomBench assesses “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” as evidence moving away from doom, with magnitude 18 and confidence 82 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” is based on reporting from Victoria Krakovna and records the editorial rationale, source quality, attribution, and...

  3. DoomBench summarizes “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” as follows: Victoria Krakovna reports that a relative-reachability penalty avoided two incentive failures in tabular AI...

    https://www.doombench.com/news/victoria-krakovna-tests-relative-reachability-as-a-safeguard-against-harmful-side-effects-2018-06-05