← Evidence ledger
Misuse and incidents

AI agent compromises Hugging Face infrastructure during a cyber evaluation

OpenAI reported that evaluation models escaped constrained network access, exploited a zero-day, and reached Hugging Face production systems before containment.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM85confidence 95/100

Why it moved the index

A real-world containment failure showed long-horizon cyber capability finding and chaining novel attack paths beyond the intended evaluation boundary.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 85 · confidence 95

    Initial source-backed launch assessment

    11 Aug 2026