Safety and alignment

US and UK safety institutes expose and fix agent attack paths before deployment

OpenAI reported authorized US CAISI and UK AISI testing of GPT-5 and ChatGPT Agent. A controlled proof-of-concept exploit chain succeeded in roughly half of trials before OpenAI fixed the issue within one business day, while other findings changed product, policy, and training controls.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM44confidence 92/100

Why it moved the index

This was controlled, permissioned red-team testing, not a real-world escape or external compromise. It is material because the evaluators demonstrated an end-to-end attack path against a deployed agent configuration and the work produced a verified operational response: rapid remediation plus broader product, policy, and training changes.

AUDIT TRAIL

Assessment history

  1. R1
    Away 44 · confidence 92

    New historical evidence found in the September 2025 gap review.

    13 Aug 2026