护盾斯特拉尔
Shieldstral
浏览论文内容
中文总结 AI 辅助
研究提出护盾斯特拉尔这一多模态安全分类器,将内容审核设为二元问答任务,通过特定数据构建方法及评估集,让小的自适应模型在文本安全基准测试中表现出色,在多模态安全分类上达新水平。
中文摘要 AI 辅助
我们引入了护盾斯特拉尔,这是一个具有30亿参数的策略自适应多模态安全分类器,在文本安全基准测试中能与近7倍于其规模的模型相匹配或表现更优,并在多模态安全分类方面创造了新的技术水平。它将内容审核制定为二元问答任务,统一多种审核任务。还介绍了数据构建方法及约5410万个样本的管理和生成,以及用于评估策略适应性的细粒度评估集,使小的自适应模型能媲美或超越更大模型。
英文摘要
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.
发表机构
- OpenAI
- Anthropic
- Google DeepMind(谷歌DeepMind)
- DeepSeek-AI(深势科技人工智能)
机构由 AI 辅助整理,请以论文原文为准。