选择性聆听:大型音频-语言模型中音频影响的机制引导控制
Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对大型音频-语言模型中无关音频干扰导致的推理偏差问题,提出ICAP-Gate机制引导控制方法,在多任务、多干扰场景下降低了配对漂移,且延迟远低于自一致性,为鲁棒多模态推理提供了新设计原则。
中文摘要 AI 辅助
大型音频-语言模型(LALMs)利用多模态证据,但在无需聆听时,与任务无关的音频会改变文本推理决策。总体准确率可能掩盖这种配对漂移,因为音频引发的修正和损害可能相互抵消。配对漂移分析和针对性干预确定了架构特定、对干预敏感的晚期音频通路作为可操作控制点。我们引入ICAP-Gate,它对每个模型的通路应用机制引导、任务条件控制。在四个LALMs、两个推理基准以及环境声和自然语音干扰下,在全部16个全拆分模型-条件评估中,ICAP-Gate的影响率和答案翻转的点估计值均低于无门控推理。固定抑制会降低所有四个模型的自动语音识别(ASR)性能,而ICAP-Gate通过保留明确音频需求指令的通路,匹配了无门控ASR性能。在全部四个评估设置中,ICAP-Gate的配对漂移点估计值均低于缓解提示,且相对于八样本自一致性(Self-Consistency)提供了有竞争力的稳定性,同时每个查询仅使用一次生成;在受控ARC测量中,自一致性的延迟是无门控的7.0至9.2倍。这些结果确立了选择性模态影响控制作为鲁棒多模态推理的设计原则。
英文摘要
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
发表机构
- College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)
- State Key Laboratory of Complex & Critical Software Environment(复杂与关键软件环境国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。