发表机构
Beihang University; Tsinghua University; State Key Laboratory of AI Safety(北京航空航天大学; 清华大学; 人工智能安全国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ThinkingGuard,一种通过逐步风险归因和结构化推理训练来检测多模态大语言模型中隐含危害的防护模型,并构建了首个风险组合性数据集TriggerBench。
AI 中文摘要
虽然多模态大语言模型(MLLMs)越来越多地部署在安全关键领域,但其可靠性受到多模态隐含风险的威胁。与显性威胁不同,这些危害是在单独良性的文本和中性视觉实体逻辑上汇聚以诱导不安全输出时出现的。当前的检测方法无法解决这一问题,因为它们忽视了控制跨模态风险激活的潜在风险激活机制,导致单模态捷径学习和幻觉合理化。为弥补这一差距,我们首先构建了TriggerBench,这是第一个明确建模风险组合性的数据集(5,600个实例)。通过正式分离关键元素和触发元素以构建反事实对比对,TriggerBench消除了风险残留,并迫使模型进行真正的逻辑演绎而非表面模式匹配,这为大规模训练和细粒度评估提供了严格的基础。在此基础上,我们提出了一种步骤监督的结构化推理训练框架,并利用它训练了ThinkingGuard,一种专门的防护模型。受态势感知理论的启发,我们将隐含风险识别解耦为渐进的认知阶段,并利用步骤奖励蒙特卡洛树搜索算法探索最优推理轨迹,然后通过双约束偏好对齐将其蒸馏到模型中。在标准安全基准和隐含安全基准上的大量实验表明,ThinkingGuard取得了强劲的性能。项目资源可在该https URL获取。
英文摘要
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
Comments9 pages, 4 figures, accepted by ACMMM 2026
Journal refIn Proceedings of the 34th ACM International Conference on Multimedia(MM '26), November 10-14, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA