Safety and alignment

Hendrycks and Mazeika publish a structured AI x-risk analysis method

The paper adapts established hazard analysis and systems-safety concepts to advanced AI, proposes X-Risk Sheets for assessing safety research, and warns that safety work can backfire when it improves general capabilities more than control.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM16confidence 68/100

Why it moved the index

The framework directly expands advanced-AI safety capacity by translating mature hazard-analysis practices into a repeatable method for identifying failure modes and checking whether proposed safety research improves the safety-to-capability balance.

AUDIT TRAIL

Assessment history

  1. R1
    Away 16 · confidence 68

    Initial inclusion from the fully reviewed, dated paper during Dan Hendrycks historical backfill.

    13 Aug 2026