arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37568cs.CL

问题中继中的魔鬼:源条件中继引导以缓解音视频大语言模型中的幻觉

Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models

Yu Zhang, Pingrui Zhang, Xuefeng Bai, Pengfei Zhang, Yang Xiang, Kehai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对音视频大语言模型的源混淆接地幻觉,通过路径干预揭示问题中继机制,提出无需训练的SECRET方法引导问题状态,在CMM和AVHBench上显著缓解幻觉。

中文摘要 AI 辅助

音视频大语言模型(AVLLMs)通过视觉、听觉和语言信息之间的交互,在多模态理解与推理方面取得了显著进展。然而,近期研究表明,AVLLMs面临一个关键挑战:源混淆的接地幻觉,即来自未使用模态的线索会引发所需模态不支持的反应,从而削弱了其在现实应用中的可靠性。现有方法在缓解这一失败方面已取得进展,但其如何源于内部跨模态交互仍未被充分理解。为弥补这一空白,我们进行了路径干预和表征分析,揭示了一种问题中继机制:问题状态携带干扰线索以及所需源证据,从而削弱了基于所需模态证据的接地。切断从干扰模态到问题状态的路径比切断到生成位置的路径能获得更大的正确答案逻辑恢复。受这些发现启发,我们提出了SECRET(源条件中继引导),一种无需训练的方法,可在问题中继处缓解跨模态干扰。通过利用不同模态路径干预引发的对比问题表征,SECRET将原始问题状态引导向所需源证据。在三个AVLLMs上对两个广泛采用的基准CMM和AVHBench进行的实验表明,SECRET始终优于先前的无训练方法,显著缓解了源混淆的接地幻觉(例如,相对于基础模型最高提升+18.0和+7.1个百分点)。模态特定字幕生成进一步证明了其对开放生成任务的泛化能力。

英文摘要

Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose $\textbf{SECRET}$ ($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.

发表机构

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Peng Cheng Laboratory(鹏城实验室)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑