Safety and alignment

OpenAI deploys free GPT-based moderation endpoint

OpenAI released a free Moderation endpoint for API-generated content, using GPT-based classifiers to identify sexual, hateful, violent, and self-harm content.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM40confidence 96/100

Why it moved the index

Free operational moderation lowered the cost of adding content controls to applications built with OpenAI models. The provider did not expose a stable exact moderation checkpoint, so no model identity is invented.

AUDIT TRAIL

Assessment history

  1. R1
    Away 40 · confidence 96

    New August 2022 deployed model-safety control absent from the durable corpus.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for OpenAI deploys free GPT-based moderation endpoint.
  1. DoomBench assesses “OpenAI deploys free GPT-based moderation endpoint” as evidence moving away from doom, with magnitude 40 and confidence 96 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “OpenAI deploys free GPT-based moderation endpoint” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “OpenAI deploys free GPT-based moderation endpoint” as follows: OpenAI released a free Moderation endpoint for API-generated content, using GPT-based classifiers to identify sexual, hateful, violent, and self-harm...

    https://www.doombench.com/news/openai-deploys-free-gpt-based-moderation-endpoint-2022-08-10