arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27548cs.AI

Nemotron 3.5 内容安全审核器:一款紧凑的多模态、多语言且具备推理能力的内容安全审核器

Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Varun Singh, Anuj Doshi, Makesh Narsimhan Sreedhar, Shaona Ghosh, Katherine Luna

中文总结 AI 辅助

本研究提出紧凑多模态多语言的Nemotron 3.5 CS安全审核器,兼具低计算成本与推理能力,可覆盖多场景安全审核,还发布配套安全数据集,验证其在多维度评估中表现优异。

中文摘要 AI 辅助

已部署AI应用的安全审核正从仅针对文本提示的方向发展:系统日益需要根据不同领域的政策对图像、文档、截图及生成的响应进行判断。现有的安全护栏通常仅覆盖该场景的部分内容,难以兼顾广泛覆盖范围、自定义政策控制与低计算成本。我们提出Nemotron 3.5内容安全审核器,本文中为简洁起见也称为Nemotron 3.5 CS,这是一款紧凑的40亿参数视觉语言安全审核器,可对12种语言的用户提示、图像及助手响应进行联合分类。Nemotron 3.5 CS可为延迟敏感型审核返回安全标签,还能生成简洁的推理轨迹,在被要求推理时应用提供的自定义政策并识别违规类别。我们还发布了用于护栏训练的多模态多语言安全数据集,涵盖人工标注的真实图像审核、良性视觉语言与文档任务、合成稀有风险及越狱案例,以及自定义政策示例。在涵盖多模态安全、文本审核、多语言鲁棒性、自定义政策遵循、良性误报及延迟的评估中,Nemotron 3.5 CS展现出实用的覆盖权衡:它新增了图像条件与政策条件的审核能力,同时仍与专用护栏模型具有广泛竞争力。这些结果表明,紧凑的视觉语言审核器可作为可部署的一线安全组件,推理能力则选择性用于审计与政策审查。

英文摘要

Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this setting, making it difficult to combine broad coverage, custom policy control, and low compute cost. We present Nemotron 3.5 Content Safety Moderator, also referred to as Nemotron 3.5 CS in this paper for brevity, a compact 4B vision-language safety moderator that jointly classifies user prompts, images, and assistant responses across 12 languages. Nemotron 3.5 CS returns safety labels for latency-sensitive moderation and can additionally produce concise reasoning traces that apply supplied custom policies and identify violated categories when reasoning is requested. We also release a multimodal and multilingual safety dataset for guard training, spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples. Across evaluations spanning multimodal safety, text moderation, multilingual robustness, custom-policy following, benign false positives, and latency, Nemotron 3.5 CS demonstrates a practical coverage tradeoff: it adds image-conditioned and policy-conditioned moderation while remaining broadly competitive with specialized guard models. These results suggest that compact vision-language moderators can serve as deployable front-line safety components, with reasoning used selectively for audit and policy review.

↑