Capability gains

OpenAI cannot rule out critical cyber capability in Astra evaluation

OpenAI said preliminary internal evaluations of its unreleased Astra model showed enough agentic coding and cyber performance that it could not rule out its Critical threshold, prompting stricter isolation, universal risky-action monitoring, and pauses on work that lacked upgraded controls.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM62confidence 55/100

Why it moved the index

A possible critical cyber threshold would directly expand consequential autonomous attack capability. Magnitude 62 reflects the zero-day and end-to-end hardened-target stakes, while confidence 55 is capped because OpenAI's evidence is preliminary, internal, and says only that Critical capability cannot yet be ruled out.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 62 · confidence 55

    Initial inclusion from OpenAI's dated disclosure of a completed preliminary capability assessment and resulting control changes.

    11 Aug 2026