Safety and alignment

David Krueger warns scalable oversight can create circular safety arguments

David Krueger argued that scalable oversight can become circular when a supposedly safe AI assistant is used to validate a stronger system before the assistant's own safety has been established. He favored robustly aligning simpler, human-judgeable behavior and using both unaided and AI-assisted human review as safety filters.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM18confidence 72/100

Why it moved the index

The post identifies a direct advanced-AI control weakness: a safety argument can depend recursively on an assistant whose own alignment has not been established. Magnitude is limited because this is conceptual analysis rather than a measured failure; confidence reflects the exact first-person reasoning and explicit date while preserving uncertainty about how broadly the mechanism applies.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 18 · confidence 72

    Adds a previously absent, dated first-person control argument found during the bounded David Krueger historical review.

    12 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for David Krueger warns scalable oversight can create circular safety arguments.
  1. DoomBench assesses “David Krueger warns scalable oversight can create circular safety arguments” as evidence moving toward doom, with magnitude 18 and confidence 72 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “David Krueger warns scalable oversight can create circular safety arguments” is based on reporting from AI Alignment Forum and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “David Krueger warns scalable oversight can create circular safety arguments” as follows: David Krueger argued that scalable oversight can become circular when a supposedly safe AI assistant is used to validate a...

    https://www.doombench.com/news/david-krueger-warns-scalable-oversight-can-create-circular-safety-arguments-2023-01-25