OpenAI · gpt-oss-safeguard

gpt-oss-safeguard-120b

READER SUMMARY

A 120B open-weight policy-reasoning model for configurable content classification and safety enforcement.

DOOM SCORE43.4out of 100model risk profile, not the overall index
CURRENT ASSESSMENT · REVISION 1

Why this model scores 43.4

Focused rather than frontier capability and little independent autonomy keep risk moderate; Apache licensing maximizes deployment while safeguard purpose limits misuse.

Capability45
Autonomy20
Deployment99
Misuse potential32
Control difficulty38
MODEL-ATTRIBUTED EVIDENCE

News tied to gpt-oss-safeguard-120b

The model score of 43.4 rates this model's risk profile. The overall Doom Index of 61.2 measures the recalibrated complete evidence corpus. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION-0.09

Each article's leave-one-out Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores are recalculated from their technical profile and related news, but never feed back into the overall index.

AUDIT TRAIL

Model score history

  1. R1
    Doom Score 43.4

    New exact safeguard model absent from the durable catalogue.

    11 Aug 2026