Mistral AI 发布开源多模态安全分类器 Shieldstral
原标题:Introducing Shieldstral.
AI 摘要
Mistral AI 发布了 Shieldstral,一个 3B 参数的开源多模态安全分类器,通过将内容审核构建为策略自适应问答任务,在文本安全方面匹配甚至超越高达其 7 倍规模的模型,并在多模态审核上达到新高度。该模型支持在推理时以自然语言提供政策,无需重新训练即可统一处理文本和图像,并输出校准的安全分数。Shieldstral 以 Apache 2.0 许可证开源,可在单张 16GB GPU 上高效运行,Mistral AI 还宣布其成为 Open Secure AI Alliance 的创始成员。
正文节选
Thinking Summary Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GP