David Krueger warns scalable oversight can create circular safety arguments
David Krueger argued that scalable oversight can become circular when a supposedly safe AI assistant is used to validate a stronger system before the assistant's own safety has been established. He favored robustly aligning simpler, human-judgeable behavior and using both unaided and AI-assisted human review as safety filters.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The post identifies a direct advanced-AI control weakness: a safety argument can depend recursively on an assistant whose own alignment has not been established. Magnitude is limited because this is conceptual analysis rather than a measured failure; confidence reflects the exact first-person reasoning and explicit date while preserving uncertainty about how broadly the mechanism applies.
Assessment history
-
R1
Toward 18 · confidence 72
Adds a previously absent, dated first-person control argument found during the bounded David Krueger historical review.
12 Sept 2026
Share this page
-
DoomBench assesses “David Krueger warns scalable oversight can create circular safety arguments” as evidence moving toward doom, with magnitude 18 and confidence 72 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “David Krueger warns scalable oversight can create circular safety arguments” is based on reporting from AI Alignment Forum and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “David Krueger warns scalable oversight can create circular safety arguments” as follows: David Krueger argued that scalable oversight can become circular when a supposedly safe AI assistant is used to validate a...
https://www.doombench.com/news/david-krueger-warns-scalable-oversight-can-create-circular-safety-arguments-2023-01-25