OpenAI uses GPT-Red to harden GPT-5.6 against prompt injection
OpenAI published GPT-Red, a self-improving automated red-team model, and separately documented its use in evaluating and training deployed GPT-5.6 safeguards against direct and agentic prompt injection.
AWAY FROM DOOM48confidence 88/100
Why it moved the index
The original dated research describes a scalable red-team model that found prompt-injection weaknesses and was used to adversarially train GPT-5.6. OpenAI's separately dated GPT-5.6 system-card update documents operational evaluation use, providing practical downstream impact rather than benchmark-only promise. The evidence is first-party, although robustness remains attack-specific rather than a general safety guarantee.
AUDIT TRAIL
Assessment history
- R1Away 48 · confidence 88
New historical safety research with separately verified practical use in deployed GPT-5.6 evaluation and training.
11 Aug 2026