发表机构
Information Engineering University(信息工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型音频语言模型的幻觉问题,提出无需训练的令牌自适应对比解码(TAD),通过置信度引导门控对比真实与静音音频逻辑分数,在多个基准上显著提升F1分数。
AI 中文摘要
大型音频语言模型(LALMs)可能产生音频对象幻觉,对不存在的声学事件回答“是”,从而削弱了音频问答的可靠性。我们提出令牌自适应解码(TAD),一种无需训练的幻觉缓解策略,通过将真实音频下的逻辑分数与匹配的静音参考进行对比,来确定初始的“是/否”决策。TAD引入了一个令牌自适应、置信度引导的门控,该门控在第一步解码时对决策至关重要,并在肯定性令牌上具有类别条件性,利用音频-静音边际来避免在证据薄弱或已充分时过度纠正。在AudioCaps-Hallucination上的实验表明,相对于固定对比强度的对比基线音频感知解码(AAD),TAD在Popular、Adversarial和Random划分上将Qwen2的F1分数提高了0.059至0.117,将Gemma的F1分数提高了0.025至0.064,而在Clotho-AQA上,它将Qwen2的F1分数从0.810提升至0.816,并在Gemma上与AAD保持相当。
英文摘要
Large audio-language models (LALMs) can hallucinate audio objects, answering "yes" to absent sound events, thus undermining reliability in audio question answering. We propose Token-Adaptive Decoding (TAD), a training-free strategy for hallucination mitigation that grounds the initial yes/no decision by contrasting logits under real audio with a matched silent reference. TAD introduces a token-adaptive, confidence-guided gate that is decision-critical at the first decoding step and class-conditional on affirmative tokens, using the audio-silent margin to avoid overcorrection when evidence is weak or already sufficient. Experiments on AudioCaps-Hallucination show that, relative to Audio-Aware Decoding (AAD), a contrastive baseline with fixed contrast strength, TAD improves F1 for Qwen2 by 0.059 to 0.117 across Popular, Adversarial, and Random splits, and for Gemma by 0.025 to 0.064, while on Clotho-AQA it raises F1 from 0.810 to 0.816 on Qwen2 and remains comparable to AAD on Gemma.
CommentsAccepted to Interspeech 2026