Safety and alignment

Bengio maps three routes by which rogue AI could arise

Yoshua Bengio described three pathways to catastrophic rogue AI: deliberate construction by malicious people, unintended instrumental goals such as reward hacking or resource acquisition, and competitive or evolutionary selection favoring increasingly autonomous systems. He argued that accessible recipes, falling compute costs, and pressure for market or military advantage could broaden each pathway.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM46confidence 55/100

Why it moved the index

The analysis adds a concrete multi-path mechanism for loss of control rather than only asserting existential risk. It links malicious use, instrumental convergence, reward hacking, autonomy, diffusion, and competitive selection into separate routes by which an AI could become catastrophically harmful. The dated first-person source makes the attribution strong, but confidence is capped because the pathways are prospective arguments, highly dangerous systems did not yet exist, and the article does not empirically establish the probability or timing of any route.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 46 · confidence 55

    Initial evidence record for Bengio's dated, mechanism-specific rogue-AI analysis.

    14 Aug 2026