arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25049cs.CLcs.AI

通过动态语义路由校准缓解大语言模型的过度拒绝问题

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration

  • School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
  • CSIRO(澳大利亚联邦科学与工业研究组织)
  • Center of Excellence for Generative AI, King Abdullah University of Science and Technology(阿卜杜拉国王科技大学生成式人工智能卓越中心)

机构由 AI 辅助整理,请以论文原文为准。

Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo

AI总结:

本文提出语义路由校准(SRC),一种轻量级、无需训练的推理框架,通过动态抑制超敏感安全头并融合双分支逻辑,缓解大语言模型对良性安全相关指令的过度拒绝,同时保持安全性能。

AI中文摘要:

为安全而对齐的大语言模型(LLMs)常常遭受过度拒绝的问题,即错误地拒绝良性但涉及安全的指令。先前的研究主要将此归因于静态表示重叠,在很大程度上忽视了潜在的动态机制。在本文中,我们通过变压器注意力内部的动态路由冲突视角,对过度拒绝进行了机制性分析。我们发现,一个稀疏的“超敏感安全头”子集在“硬安全”提示上会误触发,表现出异常的注意力纠缠,将无害的目标实体强行绑定到拒绝语义上。这引发了一个严重的、高熵的路由冲突,剥夺了目标实体所需的注意力。为应对这一问题,我们提出了语义路由校准(SRC),一个轻量级、无需训练的推理框架。SRC在推理阶段精确定位并动态抑制这些超敏感的安全头。结合作为后续解码过程中安全正则化器的双分支逻辑融合,SRC无缝地恢复了可信的推理。大量实验表明,SRC缓解了过度拒绝,同时尽可能保留了内在的安全性能。

英文摘要:

Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities to refusal semantics. This triggers a severe, high-entropy routing conflict that deprives target entities of necessary attention. To counteract this, we propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework. SRC precisely localizes and dynamically suppresses these hypersensitive safety heads at the inference stage. Coupled with a dual-branch logits fusion that acts as a safety regularizer during subsequent decoding, SRC seamlessly restores trustworthy reasoning. Extensive experiments demonstrate that SRC alleviates over-refusal, with intrinsic safety performance preserved as much as feasible.

补充信息

↑