Mistral AI has launched Shieldstral 1.0 3B, an open-weights safety classifier designed for policy adaptability. Unlike traditional models that rely on fixed harm categories, Shieldstral uses natural language questions to perform content moderation, allowing operators to define custom policies at inference time. This approach enables the model to achieve strong performance on both text and multimodal safety benchmarks, matching larger models while running efficiently on a single GPU with 16GB of VRAM. AI
IMPACT Enables flexible, on-premise content moderation for diverse applications, potentially reducing reliance on third-party vendors.
RANK_REASON Frontier-lab model release with system card.
Read on Mastodon — mastodon.social →
- Mistral AI
- Shieldstral
- Axolotl
- GPT-OSS-Safeguard-20B
- llama.cpp
- Ministral-3-3B-Base-2512
- Pixtral
- Shieldstral 1.0 3B
- vLLM
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →