Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Article URL: https://mistral.ai/news/shieldstral/ Comments URL: https://news.ycombinator.com/item?id=49171268 Points: 297 # Comments: 71
Mistral unveils Shieldstral, a 3B open-weights multimodal safety classifier that rivals models up to seven times its size. This model transforms content moderation into a policy-adaptive question-answering task, allowing plain-language policies to be applied at inference time. Unlike traditional guardrail models, Shieldstral accepts policies in natural language, unifying text and image safety evaluation without requiring retraining.
Available under the Apache 2.0 license, Shieldstral delivers calibrated safety scores across various benchmarks while running efficiently on a single 16GB NVIDIA GPU. The model's adaptability enables it to handle diverse content types, including text, images, and prompt-response pairs, with a single interface. Shieldstral's performance surpasses open guard models of similar size in text safety, refusal detection, policy adaptability, and multimodal benchmarks. trained on a diverse set of real and synthetic data, the model learns to distinguish specific policies from similar, easily confused ones, making it highly effective at adapting to novel policies at deployment time.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.