Safety and alignment

Mistral releases open Shieldstral safety classifier

Mistral released Shieldstral 1.0 3B under Apache 2.0 as a policy-adaptive text and image safety classifier, with held-out benchmark results and operation on a single 16 GB GPU.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 80/100

Why it moved the index

An openly available, policy-adaptive moderation model makes a concrete safeguard easier to deploy, although benchmark performance does not establish its effectiveness in consequential real-world use.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 80

    New dated primary release and model card establish the exact open safety model, availability, benchmark scope, and deployment requirements.

    13 Aug 2026