Safety and alignment

OpenAI deploys InstructGPT as default API language models

OpenAI made InstructGPT models the default language models on its API after human-feedback training improved instruction following and reduced toxic, untruthful, and hallucinatory outputs.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM37confidence 95/100

Why it moved the index

Deploying human-feedback-trained models as the API default translated alignment research into broad practical safeguards and reduced several measured harmful behaviors. Residual failures and the lack of a disclosed exact serving checkpoint constrain the magnitude and preclude a model record.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 37 · confidence 95

    New January 2022 deployed alignment intervention absent from the durable corpus; exact serving checkpoint remains intentionally unprofiled.

    12 Aug 2026