Safety and alignment

OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks

Automated red teaming found new long-horizon prompt-injection attacks, leading OpenAI to deploy a hardened browser-agent checkpoint and strengthened safeguards to all Atlas users.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 74/100

Why it moved the index

The work produced a concrete production checkpoint and layered defenses against realistic agent hijacking, although OpenAI explicitly states prompt injection remains an open challenge without deterministic guarantees.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 74

    New dated production safeguard deployment absent from durable context.

    11 Aug 2026