arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04004cs.SD

基于多模态直接偏好优化减少大语言模型音频理解中的幻觉

Hallucination Reduction for LLM-Based Audio Understanding via Multimodal Direct Preference Optimization

Bebe Cosgrove, Aaron Isidore Grace, Weiran Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型音频语言模型在音频与文本输入下产生幻觉的问题,提出多模态直接偏好优化方法,通过对比完整与失真音频迫使模型基于声学输入,在DCASE 2025和AH数据集上分别提升14.0%和27.4%的准确率。

中文摘要 AI 辅助

大型音频语言模型(LALMs)在同时接收音频和文本输入时,容易产生幻觉并过度依赖文本先验。为缓解这些幻觉,我们提出利用多模态直接偏好优化(mDPO)目标,通过对比完整与失真的音频对应物,迫使模型将其生成基于声学输入。我们将这一偏好学习框架扩展到音频领域,应用多种声学扰动。在DCASE 2025挑战赛和AH存在性数据集上评估Qwen2-Audio骨干模型,我们证明将mDPO扩展到LALMs显著增强了复杂DCASE数据集中的时间推理能力,并在AH基准上提升了基本存在性验证的性能。我们确定时间反转、频率掩蔽和随机噪声为最有效的扰动。最终,我们的方法在DCASE 2025数据集上实现了14.0%的绝对准确率提升,在AH存在性上实现了27.4%的提升。

英文摘要

Large Audio Language Models (LALMs) are prone to hallucinating and over-relying on text priors when simultaneously presented with audio and text inputs. To mitigate these hallucinations, we propose utilizing the multimodal Direct Preference Optimization (mDPO) objective, which forces the model to ground its generation in the acoustic input by contrasting intact and distorted audio counterparts. We extend this preference learning framework to the audio domain by applying a variety of acoustic perturbations. Evaluating the Qwen2-Audio backbone across the DCASE 2025 Challenge and AH Existence datasets, we demonstrate that extending mDPO to LALMs significantly enhances temporal reasoning in the complex DCASE dataset, and improves performance on basic existence verification in the AH benchmark. We identify temporal reversal, frequency masking, and random noise as the most effective perturbations. Ultimately, our approach achieves an absolute accuracy improvement of 14.0% on the DCASE 2025 dataset and 27.4% on AH Existence.

发表机构

  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑