arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07774cs.CL

读取而非操纵:利用路由器对数提升MoE视觉语言模型的多模态安全性

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出利用MoE视觉语言模型路由器对数作为多模态安全诊断信号,开发轻量级检测器在生成前识别不安全请求,无需修改模型参数,显著降低安全错误并良好泛化至分布外基准。

中文摘要 AI 辅助

视觉语言模型(VLM)面临组合性安全风险,其中有害意图源于视觉与文本输入之间的交互。随着混合专家(MoE)视觉语言模型日益普遍,近期工作探索了多种安全干预措施,包括提示、监督微调和基于路由的专家引导。然而,这些方法在模型和评估分布上的改进不一致,且对模型行为或内部状态的干预会因过度拒绝而引入安全-效用权衡。我们不操纵内部状态来引导模型行为,而是探究路由状态能否作为多模态安全性的诊断信号。我们发现,路由器对数确实能对多模态输入是否安全提供高度预测性的信号。基于这一观察,我们引入了一个轻量级的路由器对数安全检测器,它在提示预填充期间读取路由信号,并在生成之前识别不安全请求,而无需修改模型参数或专家路由。在Qwen3-VL和Kimi-VL上,所提出的检测器显著减少了HoliSafe基准上的安全错误,并出色地泛化到具有不同安全模式的分布外安全基准,包括MISHard和MM-SafetyBench。所提出的路由器对数检测器的成功也暗示了对模型内部结构的更广阔视角:与其仅关注操纵内部组件以引导行为,简单地读取自然出现的信号并将其链接到外部安全机制,可以为现有安全干预提供一种简单、有效且非侵入性的补充。

英文摘要

Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.

发表机构

  • University of Washington(华盛顿大学)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑