Safety and alignment

OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks

Automated red teaming found new long-horizon prompt-injection attacks, leading OpenAI to deploy a hardened browser-agent checkpoint and strengthened safeguards to all Atlas users.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 74/100

Why it moved the index

The work produced a concrete production checkpoint and layered defenses against realistic agent hijacking, although OpenAI explicitly states prompt injection remains an open challenge without deterministic guarantees.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 74

    New dated production safeguard deployment absent from durable context.

    11 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks.
  1. DoomBench assesses “OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks” as evidence moving away from doom, with magnitude 35 and confidence 74 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision...

  3. DoomBench summarizes “OpenAI ships an adversarially trained Atlas checkpoint after new prompt-injection attacks” as follows: Automated red teaming found new long-horizon prompt-injection attacks, leading OpenAI to deploy a hardened...

    https://www.doombench.com/news/openai-ships-an-adversarially-trained-atlas-checkpoint-after-new-prompt-injection-attacks-2025-12-22