arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

音频语言模型中拒绝方向的原因定位

Causal Localization of the Refusal Direction in Audio Language Models

Leonardo Haw-Yang Foo, Hung-yi Lee

arXiv 2609.22260首次发表:更新:

发表机构

National Taiwan University; NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(台湾大学; 台湾大学人工智能卓越研究中心(台大AI卓越研究中心))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过因果干预定位大型音频语言模型中拒绝边际的依赖层,发现主要效应在文本LM中后期层,而非音频接口,并建议安全审计采用干预方法。

AI 中文摘要

大型音频语言模型(LALM)将语音前端附加到已经经过安全对齐的文本语言模型(LM)上。当这样的模型拒绝有害的语音请求时,拒绝是由前端承载的,还是从文本LM继承的?我们通过因果干预来测试这一点。在每个模型的音频到LM接口以及测试的LM残差层上,我们拟合一个区分有害提示与良性提示的方向,消融其分量,并测量由此产生的模型首词拒绝边际的变化。五个模型中有四个在保留类别转移下进行评估。跨越三个骨干家族的五个LALM,其中三个模型通过了基线安全门,测试到的最大效应出现在中到后期LM频带,而测试接口方向的消融影响很小。在Qwen2.5-Omni上,消融L16方向使边际变化-7.10,而在投影仪处为-0.013。音频通路仍在使用:将编码器输出置零会使边际变化-4.7。在同一模型中,对比在早期层可线性解码,而单层消融影响很小,且仅在LM骨干上拟合的方向可转移到完整音频模型。由于有害和良性提示在形式上也有所不同,我们将该方向解释为与拒绝相关而非特定于有害性。这些干预定位了拒绝边际的依赖性,而非拒绝计算的位置。移动边际也不总是改变模型所写的内容。对这些模型的安全审计应使用干预而非仅依赖探针,并应检查继承的文本LM以及音频接口。

英文摘要

A large audio language model (LALM) attaches a speech front end to a text language model (LM) that is already safety-aligned. When such a model refuses a harmful spoken request, is the refusal carried by the front end, or inherited from the text LM? We test this with causal interventions. At each model's audio-to-LM interface and at tested LM residual layers, we fit a direction separating harmful from benign prompts, ablate its component, and measure the resulting change in the model's first-token refusal margin. Four of the five models are evaluated under held-out category shift. Across five LALMs spanning three backbone families, with three models passing a baseline safety gate, the largest tested effects occur in a mid-to-late LM band, while ablations of the tested interface directions have little effect. On Qwen2.5-Omni, ablating the L16 direction changes the margin by -7.10, versus -0.013 at the projector. The audio pathway is still in use: zeroing the encoder output changes the margin by -4.7. In the same model, the contrast is linearly decodable at an early layer where single-layer ablation has little effect, and a direction fitted on the LM backbone alone transfers to the full audio model. Because harmful and benign prompts also differ in form, we interpret the direction as refusal-linked rather than harmfulness-specific. These interventions localize dependence of the refusal margin, not where refusal is computed. Moving the margin also does not always change what the model writes. Safety audits of these models should use interventions rather than rely on probes alone, and should examine the inherited text LM alongside the audio interface.

CommentsAccepted at ISCSLP 2026. 4 pages + references

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑