Safety and alignment

Ajeya Cotra proposes aligning narrowly superhuman models as a human-oversight testbed

In a dated first-person essay, Ajeya Cotra proposed training and evaluating existing models on fuzzy tasks where they may outperform some human demonstrators or judges. She argued that studying whether weaker humans can supervise these systems could provide a practical testbed for scalable oversight of more capable AI. The article proposes research; it does not report a validated safeguard.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM18confidence 55/100

Why it moved the index

Magnitude 18 reflects a concrete research direction for the central problem of supervising models that can exceed human evaluators, without credit for implementation or demonstrated risk reduction. Confidence 55 reflects a fully attributable and precisely dated primary argument with a direct oversight nexus, but its transfer from narrow testbeds to future advanced systems remains an unvalidated inference. GPT-3 is an illustrative example rather than an assessed exact-model development.

AUDIT TRAIL

Assessment history

  1. R1
    Away 18 · confidence 55

    Distinct original 2021 scalable-oversight proposal, absent from fresh durable context.

    23 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Ajeya Cotra proposes aligning narrowly superhuman models as a human-oversight testbed.
  1. DoomBench assesses “Ajeya Cotra proposes aligning narrowly superhuman models as a human-oversight testbed” as evidence moving away from doom, with magnitude 18 and confidence 55 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Ajeya Cotra proposes aligning narrowly superhuman models as a human-oversight testbed” is based on reporting from AI Alignment Forum and records the editorial rationale, source quality, attribution, and...

  3. DoomBench summarizes “Ajeya Cotra proposes aligning narrowly superhuman models as a human-oversight testbed” as follows: In a dated first-person essay, Ajeya Cotra proposed training and evaluating existing models on fuzzy tasks where...

    https://www.doombench.com/news/ajeya-cotra-proposes-aligning-narrowly-superhuman-models-as-a-human-oversight-testbed-2021-03-05