Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-answering: a plain-language policy question is supplied at inference time and the model's yes/no token logits are normalized into a continuous safety score, letting one model serve divergent moderation taxonomies without retraining. Covers text, image, and text+image inputs.
Key Findings
- Training data consolidates approximately 54.1M samples from heterogeneous safety datasets with divergent taxonomies under the single binary question-answering formulation
- Reported to match or outperform guard models nearly 7x its size on text safety benchmarks and to set a new state of the art on multimodal safety classification
- Ships as open weights (Apache 2.0) with a fine-grained evaluation set for policy adaptability; runs on a single 16GB GPU
Methodology Notes
arXiv technical report 2607.25857 (v1 2026-07-28, v2 2026-08-04), verified via the arXiv API; announcement, model card, and Hugging Face weights released 2026-08-04. Benchmark comparisons are Mistral's own reported numbers. Multilingual coverage is framed as future work in the announcement.
Sources
Tags
Cite This
APA
Mistral AI (2026). Shieldstral. Mistral AI. https://arxiv.org/abs/2607.25857