Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models
French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperforms other large language models up to seven times its size, setting a new standard for moderation. The new model, named Shieldstral, allows developers to write policies in natural language […] The post…
Mistral AI SAS, a French artificial intelligence company, unveiled Shieldstral, a compact and efficient multimodal safety AI model. The model demonstrates superior performance compared to other language models, even those of a larger size. Named Shieldstral, this open-weight model allows developers to define policies using natural language questions. The model then processes the content, providing a safety score and a simple "yes" or "no" verdict.
Shieldstral's primary strengths lie in its text and image safety capabilities. It surpasses other models by a significant margin, scoring an overall average of 84.9% in text safety benchmarks and 83.8% in multimodal image safety benchmarks. The model operates on a straightforward setup: developers input a high-level task, a user query, and the content.
Shieldstral's distinct capability is its ability to distinguish between closely related but different policies, enabling fine-grained content classification. For instance, it can differentiate between content related to malware instructions and cybersecurity discussions, assigning the former to a "malware instructions" policy (a violation) and the latter to a "cybersecurity discussion" policy (acceptable).
Developers can easily customize and adapt policies at runtime, making Shieldstral versatile for various applications like customer service text safety checks, AI assistant refusal detection, policy violations, and image generation security. Despite its impressive performance, Shieldstral remains lightweight, with a model size of just 3 billion parameters. This allows it to run efficiently on a single 16-gigabyte graphics processing unit and integrate swiftly with larger models for added safety guardrails.
This innovative model is part of Mistral AI's broader mission to make moderation more context-aware, natural, and adaptable, particularly in the realm of multilingual and longer-document coverage.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.