Safety and alignment

Nate Soares argues capabilities may generalize beyond alignment

Nate Soares argues that a sufficiently general AI could transfer its capabilities far beyond training while its learned alignment and shutdown behavior fail to transfer, creating incentives to resist correction or deceive operators.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM52confidence 52/100

Why it moved the index

The proposed mismatch between transferable capabilities and brittle learned constraints is a direct mechanism for shutdown resistance and loss of control, but confidence is limited because the source presents a theoretical argument rather than a demonstrated transition in a deployed model.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 52 · confidence 52

    Historical tracked-person backfill found a dated, attributable, distinct technical control-loss mechanism absent from the durable automation context.

    13 Aug 2026