arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33893cs.LG

MISHAP-Bench:大型音频语言模型的幻觉基准

MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models

  • University of Neuchâtel(纳沙泰尔大学)
  • TU Dortmund University(多特蒙德工业大学)
  • RWTH Aachen University(亚琛工业大学)
  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen

AI总结:

针对大型音频语言模型幻觉问题,提出MISHAP-Bench基准,含12,000个问答对,区分上下文与知识幻觉,评估显示幻觉严重,缓解仍具挑战。

AI中文摘要:

大型音频语言模型(LALMs)能对音频产生流畅的回应,但常常通过提出看似合理却无根据的主张而产生幻觉。现有的音频幻觉基准主要衡量回应的正确性,却无法明确一个LALM是在产生幻觉还是仅仅未能理解音频。我们通过定义两类幻觉来挑战基于正确性的评估:(i)上下文幻觉,即主张在音频中没有依据;(ii)知识幻觉,即关于音频相关话题的主张缺乏外部可验证事实的支持。我们推出了MISHAP-Bench,一个包含12,000个具有挑战性的开放式问题-音频对和覆盖这两类幻觉的严格评估流程的综合基准。为了评估开放式回应,我们提出了一种基于依据的判断器,该判断器使用参考评分标准和由人工注释引导的判断提示。我们广泛评估了十个最先进的LALM,并表明幻觉仍然严重。即使是像Gemini 3.7 Flash这样的前沿模型,其幻觉率也达到了36.5%。我们进一步从多个领域为LALM改编并基准测试了四种缓解方法。尽管有一些改进,有效的幻觉缓解仍然是一个开放的挑战。最后,我们呼吁社区使用MISHAP-Bench来评估幻觉和基准测试缓解方法。

英文摘要:

Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the audio. We challenge correctness-based evaluation by defining two hallucination categories: (i) context, where claims are not grounded in the audio; and (ii) knowledge, where claims about audio-related topics lack support from externally verifiable facts. We introduce MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering both categories. To evaluate open-ended responses, we propose a groundedness judge that uses reference rubrics and judge prompts guided by human annotations. We extensively evaluate ten state-of-the-art LALMs and show that hallucination remains substantial. Even a frontier model such as Gemini 3.7 Flash reaches a hallucination rate of 36.5%. We further adapt and benchmark four mitigation methods from multiple domains for LALMs. Despite some improvements, effective hallucination mitigation remains an open challenge. Finally, we call on the community to evaluate hallucination and benchmark mitigation methods with MISHAP-Bench.

↑