arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09697cs.LG

安全响应很重要:输出感知安全护栏减轻多模态大语言模型中的过度拒绝

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs

Jiayi Li, Kun Zhan

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态大语言模型安全机制权衡问题,提出输出感知安全护栏范式,通过在模型隐藏状态空间预测及多实例对比学习训练分类器,减少过度拒绝,匹配安全性能,保留模型效用与安全能力。

中文摘要 AI 辅助

现有多模态大语言模型(MLLMs)的安全机制在安全与效用之间面临根本权衡。模型微调实现了强大的安全性,但损害了通用效用。输入侧安全护栏提供了一种轻量级替代方案,但存在严重的过度拒绝问题,会不加区分地阻止良性查询或模型本可通过拒绝或建议性回复安全回答的查询。我们发现过度拒绝的根本原因在于输入感知范式,安全护栏在不考虑模型自身是否能够生成安全响应的情况下做出安全决策。通常,MLLMs已经拥有内在安全机制,可将有害输入转化为无害输出,但输入侧安全护栏会覆盖此能力,降低用户体验。基于此,我们提出向输出感知安全护栏的范式转变。我们的方法在模型的隐藏状态空间内运行,在即将生成的内容完全生成之前预测其是否不安全。通过对隐藏状态表示进行多实例对比学习来训练轻量级分类器,我们的方法能够区分会导致不安全输出的输入和不会导致不安全输出的输入,即使输入本身包含风险元素。这使得仅在模型的实际响应会有害时进行精确干预。大量实验表明,我们的输出感知安全护栏在大幅减少过度拒绝的同时,与现有方法的安全性能相匹配,保留了模型的效用和内置安全能力。

英文摘要

Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility. Model fine-tuning achieves robust safety but compromises general utility. Input-side safety guardrails offer a lightweight alternative, yet they suffer from severe over-refusal, indiscriminately blocking benign queries or those the model could have safely answered through refusal or advisory responses. We identify that the root cause of over-refusal lies in the input-aware paradigm: safety guardrails make safety decisions without considering whether the model itself is capable of generating safe responses. Usually, MLLMs already possess intrinsic safety mechanisms that can transform harmful inputs into harmless outputs, but input-side safety guardrails override this capability, degrading user experience. Motivated by this insight, we propose a paradigm shift toward output-aware safety guardrails. Our method operates within the model's hidden state space to predict whether the forthcoming generation will be unsafe before it is fully produced. By training a lightweight classifier via multi-instance contrastive learning on hidden state representations, our approach distinguishes between inputs that will lead to unsafe outputs and those that will not, even when the inputs themselves contain risky elements. This enables precise intervention only when the model's actual response would be harmful. Extensive experiments demonstrate that our output-aware safety guardrail matches the safety performance of existing methods while drastically reducing over-refusal, preserving the model's utility and built-in safety capabilities. Code is available at: https://github.com/kunzhan/OutGuard

发表机构

  • Lanzhou University(兰州大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑