all the models — AI benchmark observatory
← Models

Shieldstral 1.0 (3B)

Mistral AI · open weight · shieldstral-1.0-3b

Shieldstral is a 3B open-weight multimodal safety classifier from Mistral. It frames content moderation as policy-adaptive yes/no question answering: plain-language policies are supplied at inference time, and the model returns a calibrated safety score from a single forward pass for text, image, or text+image content. Released under Apache 2.0. Aggregate self-reported average F1: text-safety 0.849, multimodal 0.838 (HF mistralai/Shieldstral-1.0-3B / arXiv 2607.25857); per-bench WildGuard/HarmBench/ToxicChat/BeaverTails/Aegis/VLGuard ids are not in the catalog.

Not enough scores for a fingerprint yet.

Benchmark scores

BenchmarkScore

No scores synced for this model.

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.