发表机构
Sungkyunkwan University; Daegu Gyeongbuk Institute of Science and Technology (DGIST)(成均馆大学; 大邱庆北科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无需训练的RED方法,通过点互信息识别问题相关证据并增强其贡献,在三个音频-视觉幻觉基准上显著提升准确率,缓解跨模态幻觉。
AI 中文摘要
音频-视觉大语言模型(AV-LLMs)仍然容易产生跨模态幻觉,即一种模态错误地影响对另一种模态的预测。尽管对比解码减少了视觉-语言模型中的幻觉,但其直接扩展到AV-LLMs时忽略了一个关键挑战:不同的问题需要不同的感知证据,包括音频、视频或它们的交互。值得注意的是,我们观察到,即使模型能够从单一信息模态中恢复正确答案,联合音频-视觉推理也可能削弱预测。例如,当被问及听到什么乐器时,模型可能仅从音频中正确预测小提琴。一旦添加显示吉他的视频,其对小提琴的置信度可能会下降。在本文中,我们引入了相关证据解码(RED),一种无需训练的方法,用于识别与问题相关的证据并选择性地加强其贡献。RED使用点互信息来量化音频和视频在问题之外提供的预测支持。它将它们的联合贡献分解为音频、视频和残差交互组件。一次仅问题的推理过程确定所需的证据类型,之后模型用相应的PMI贡献增强原始音频-视觉预测。在三个音频-视觉幻觉基准和三个AV-LLMs上,RED相比标准解码将CMM上的准确率提高了高达7%,AVHBench上提高了6.3%,SVHalluc上提高了3.8%,平均相对首令牌时间为标准解码的1.5倍。
英文摘要
Audio-Visual Large Language Models (AV-LLMs) remain prone to cross-modal hallucinations, where one modality incorrectly affects predictions about another. Although contrastive decoding reduces hallucinations in vision-language models, its direct extension to AV-LLMs overlooks a key challenge: different questions require different perceptual evidence, including audio, video, or their interaction. Notably, we observe that joint audio-visual inference can weaken the prediction even when a model can recover the correct answer from a single informative modality. For example, when asked which instrument is heard, a model may correctly predict violin from the audio alone. Once a video showing a guitar is added, its confidence in violin may drop. In this paper, we introduce Relevant Evidence Decoding (RED), a training-free method that identifies question-relevant evidence and selectively strengthens its contribution. RED uses pointwise mutual information to quantify the predictive support provided by audio and video beyond the question alone. It decomposes their joint contribution into audio, video, and residual interaction components. A question-only inference pass determines the required evidence type, after which the model augments the original audio-visual prediction with the corresponding PMI contribution. Across three audio-visual hallucination benchmarks and three AV-LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, with an average relative time to first token of 1.5x standard decoding.