Victoria Krakovna tests relative reachability as a safeguard against harmful side effects
Victoria Krakovna reports that a relative-reachability penalty avoided two incentive failures in tabular AI Safety Gridworlds: blocking irreversible events and undoing helpful interventions. The proof of concept preserved safe options better than reversibility and simple impact penalties, while leaving realistic scale and default-outcome definitions unresolved.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The first-person report documents a concrete safeguard that avoided specific side-effect incentives in two controlled gridworlds. Its direct human-control relevance is clear, but the result is small-scale and the author explicitly leaves realistic tractability, default-outcome definition, and omitted-preference failures unresolved.
Assessment history
-
R1
Away 18 · confidence 82
Adds a previously absent, exactly timestamped proof of concept for reducing harmful AI side effects and its stated limitations.
22 Sept 2026
Share this page
-
DoomBench assesses “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” as evidence moving away from doom, with magnitude 18 and confidence 82 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” is based on reporting from Victoria Krakovna and records the editorial rationale, source quality, attribution, and...
-
DoomBench summarizes “Victoria Krakovna tests relative reachability as a safeguard against harmful side effects” as follows: Victoria Krakovna reports that a relative-reachability penalty avoided two incentive failures in tabular AI...
https://www.doombench.com/news/victoria-krakovna-tests-relative-reachability-as-a-safeguard-against-harmful-side-effects-2018-06-05