Skip to main content
Lab publication Credible

Shieldstral

Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-answering: a plain-language policy question is supplied at inference time and the model's yes/no token logits are normalized into a continuous safety score, letting one model serve divergent moderation taxonomies without retraining. Covers text, image, and text+image inputs.

Publisher

Mistral AI

Published

28 Jul 2026

Added

2 weeks ago

Key Findings

  • Training data consolidates approximately 54.1M samples from heterogeneous safety datasets with divergent taxonomies under the single binary question-answering formulation
  • Reported to match or outperform guard models nearly 7x its size on text safety benchmarks and to set a new state of the art on multimodal safety classification
  • Ships as open weights (Apache 2.0) with a fine-grained evaluation set for policy adaptability; runs on a single 16GB GPU

Methodology Notes

arXiv technical report 2607.25857 (v1 2026-07-28, v2 2026-08-04), verified via the arXiv API; announcement, model card, and Hugging Face weights released 2026-08-04. Benchmark comparisons are Mistral's own reported numbers. Multilingual coverage is framed as future work in the announcement.

Tags

mistralguard-modelopen-weightsmoderationmultimodal

Cite This

APA

Mistral AI (2026). Shieldstral. Mistral AI. https://arxiv.org/abs/2607.25857