Safety and alignment

OpenAI deploys free GPT-based moderation endpoint

OpenAI released a free Moderation endpoint for API-generated content, using GPT-based classifiers to identify sexual, hateful, violent, and self-harm content.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM40confidence 96/100

Why it moved the index

Free operational moderation lowered the cost of adding content controls to applications built with OpenAI models. The provider did not expose a stable exact moderation checkpoint, so no model identity is invented.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 40 · confidence 96

    New August 2022 deployed model-safety control absent from the durable corpus.

    12 Aug 2026