← Evidence ledger
Safety and alignment

OpenAI uses GPT-Red to harden GPT-5.6 against prompt injection

OpenAI published GPT-Red, a self-improving automated red-team model, and separately documented its use in evaluating and training deployed GPT-5.6 safeguards against direct and agentic prompt injection.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM48confidence 88/100

Why it moved the index

The original dated research describes a scalable red-team model that found prompt-injection weaknesses and was used to adversarially train GPT-5.6. OpenAI's separately dated GPT-5.6 system-card update documents operational evaluation use, providing practical downstream impact rather than benchmark-only promise. The evidence is first-party, although robustness remains attack-specific rather than a general safety guarantee.

AUDIT TRAIL

Assessment history

  1. R1
    Away 48 · confidence 88

    New historical safety research with separately verified practical use in deployed GPT-5.6 evaluation and training.

    11 Aug 2026