发表机构
Macau University of Science and Technology; Wuhan University of Technology; Nanjing Institute of Technology; National Tsing Hua University(澳门科技大学; 武汉理工大学; 南京工程学院; 国立清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大模型视觉问答中安全决策依赖表面相关性的问题,提出A$^2$Safe框架,通过反事实证据对齐与自适应智能体协作,在安全关键和通用协议下分别达到95.72安全分数和78.34通用分数,显著降低良性拒绝率。
AI 中文摘要
多模态大语言模型(MLLMs)的视觉问答(VQA)不仅需要生成安全有效的响应,还需要将安全决策建立在决定风险的多模态证据之上。近期安全对齐方法改善了拒绝行为与情境风险感知,但正确的安全结果仍可能依赖于表面的文本或视觉相关性,尤其是当风险源于单个良性的图像与问题内容之间的交互时。为解决这一问题,我们提出了A$^2$Safe,一种反事实证据对齐的自适应智能体协作框架,用于安全有效的VQA。A$^2$Safe通过接地安全证据板(Grounded Safety Evidence Board)组织局部视觉观察、文本意图和跨模态风险关系,使安全决策的依据明确化。反事实安全证据对齐强制对与安全无关的变化保持不变性,同时在风险关键证据发生最小改变时要求适当的安全状态和响应模式转换。由此产生的证据状态进一步支持自适应协作,当接地证据充分时直接回答,当证据存在风险、不确定或冲突时触发策略批评和响应修订。在互补的安全关键和通用VQA协议下,A$^2$Safe取得了95.72的SIUO安全分数,将MOSSBench上的良性拒绝率降至14.67%,并以27.8%的令牌开销维持了78.34的平均通用VQA分数。这些结果支持了反事实证据对齐的自适应协作,用于安全有效的多模态问答。
英文摘要
Visual Question Answering (VQA) with Multimodal Large Language Models (MLLMs) requires not only producing safe and effective responses, but also grounding safety decisions in the multimodal evidence that determines risk. Recent safety-alignment methods improve refusal behavior and contextual risk awareness, yet correct safety outcomes may still rely on superficial textual or visual correlations, particularly when risk emerges from interactions between individually benign image and question content. To address this issue, we propose A$^2$Safe, a counterfactual evidence-aligned adaptive agent collaboration framework for safe and effective VQA. A$^2$Safe organizes localized visual observations, textual intent, and cross-modal risk relations through a Grounded Safety Evidence Board, making the basis of safety decisions explicit. Counterfactual safety evidence alignment enforces invariance to safety-irrelevant changes while requiring appropriate safety-state and response-mode transitions when risk-critical evidence is minimally altered. The resulting evidence state further supports adaptive collaboration, enabling direct answering when grounded evidence is sufficient and invoking policy critique and response revision when evidence is risky, uncertain, or conflicting. Under complementary safety-critical and general VQA protocols, A$^2$Safe achieves a 95.72 SIUO safety score, reduces the benign refusal rate on MOSSBench to 14.67%, and maintains an average general VQA score of 78.34 with 27.8% token overhead. These results support counterfactual evidence-aligned adaptive collaboration for safe and effective multimodal question answering.