Safety and alignment

Jan Leike identifies self-exfiltration as a key AI control threshold

Jan Leike argues that a model able to copy its own weights beyond an operator's servers could become practically irrecoverable, making self-exfiltration capability a critical safety and deployment threshold.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM48confidence 58/100

Why it moved the index

Self-exfiltration would create a direct and potentially irreversible loss-of-control pathway by moving model weights beyond an operator's shutdown authority, giving the mechanism material long-run significance; confidence is capped because this is a specific threat model and evaluation proposal, not an observed incident.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 48 · confidence 58

    Historical tracked-person backfill found a dated, attributable, distinct control-loss mechanism absent from the durable automation context.

    13 Aug 2026