发表机构
Northeastern University; Bose Corporation; Stony Brook University(东北大学; 博士公司; 石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出MRMAD多轮多音频基准,评估18种LALMs的音频降级感知,发现现有模型难可靠推理降级,为构建鲁棒LALMs提供基础。
AI 中文摘要
大型音频语言模型(LALMs)在理解语音、音乐和通用声音事件方面已展现出良好进展,但其对音频信号降级的推理能力仍未得到充分探索。现有基准主要评估语义理解、事件识别或高级音频推理,留下一个基本问题未解答:LALMs是否理解音频质量的差异?我们提出MRMAD,即用于评估LALMs中音频降级感知与理解的多轮多音频降级基准。MRMAD涵盖语音、音乐和声音,将评估构建为针对多个音频输入的多轮对话,要求模型识别降级类型、比较严重程度并感知多轮中的损坏变化。与当前的单轮音频语言基准不同,MRMAD评估LALMs能否基于新证据维持一致的降级假设,并以自然语言解释低级声学现象。通过对18个具有代表性的LALMs(从非思考型到推理型及全模态模型)进行系统评估,我们发现当前模型通常能识别粗略内容,但无法可靠地诊断、比较或推理降级。MRMAD揭示了音频语言理解中一个重要但被忽视的方面,并为构建未来对现实世界声学条件具有鲁棒性的LALMs提供了诊断基础。
英文摘要
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues across multiple audio inputs, requiring models to identify types of degradation, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and comprehend low-level acoustic phenomena over multi-turn dialogues. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. Human evaluations further reveal a significant perception gap between LALMs and human listeners. MRMAD thus exposes a critical yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.
CommentsEMNLP 2026, Code and benchmark: https://github.com/Bose/MRMAD