arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29098cs.AIcs.CV

SafeAtlas-VL:基于大规模数据与防护模型的二元多模态安全之外

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei … 展开作者

Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现有多模态安全评估的二元局限,推出含五级标注的SafeAtlas-VL数据集与SafeAtlas Guard系列模型,8B模型F1分数超SOTA约4%,代码数据模型已公开。

中文摘要 AI 辅助

多模态安全 moderation(审核)需要区分视觉内容、用户意图与助手行为引发的风险。然而,现有防护措施通常针对单一判断目标进行训练,将安全评估简化为二元决策,导致多模态交互中的风险难以比较,模糊案例被掩盖。我们推出SafeAtlas-VL,这是一个包含150万训练实例的数据集,对图像、请求和响应层面的判断采用五级有序量表。我们从真实世界与合成来源中整理了广泛的安全相关数据,并应用了感知分歧的标注流程。该数据集涵盖15个伤害类别与55个细分子类别,覆盖了广泛的多模态安全场景。我们还构建了SafeAtlas-Bench,这是一个包含5000个实例的保留集,用于评估五级预测与连续风险分数。基于该数据集,我们通过目标条件调优训练了SafeAtlas Guard系列模型,用于多模态安全检测。我们的模型不仅能对安全级别进行五分类,还能通过软累积序数头将安全映射为连续分数。实验结果表明,在我们的数据集上训练的防护模型表现出强大的泛化能力:即使不使用其他基准的训练集,它们也能在相应测试集上取得有竞争力的性能。值得注意的是,我们的8B模型实现了整体最佳性能,F1分数比之前的SOTA(最优技术)高出约4%。代码、数据与模型已发布以支持进一步研究。警告:本文包含可能具有冒犯性、有害性、图形化或令人不安的示例数据。

英文摘要

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑