发表机构
Schulich School of Music, McGill University; CIRMMT (Centre for Interdisciplinary Research in Music Media and Technology); Courant Institute, New York University; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(麦吉尔大学舒利希音乐学院; 音乐媒体与技术跨学科研究中心(CIRMMT); 纽约大学柯朗研究所; 穆罕默德·本·扎耶德人工智能大学(MBZUAI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出时序音乐定位任务,构建基准集MusicGroundingBench评估音频-语言模型的该能力,发现当前模型完成该任务仍具挑战,任务特定训练可提升性能,该基准可用于评估模型回应是否基于音乐证据。
AI 中文摘要
大型音频-语言模型能够生成流畅且符合音乐逻辑的回应,但这些回应是否基于音频输入仍不明确。本文提出时序音乐定位任务,要求模型返回对应查询音乐音符、事件或模式的一个或多个时间跨度。为评估该能力,我们构建了受控基准测试集MusicGroundingBench,通过将算法生成的钢琴MIDI渲染为音频,实现精确的符号-音频对齐,包含两个子集:MGBench-3N评估最多含3个音符的片段的音符级定位,MGBench-2B评估两小节节选的结构化定位与短时长音乐理解。实验表明,时序音乐定位对当前音频-语言模型仍具挑战性,而任务特定训练可带来显著提升。我们还报告了关于定位监督与音乐理解关系的探索性证据。这些结果确立了MusicGroundingBench作为受控测试平台,用于评估音频-语言模型的回应是否基于时序定位的音乐证据。
英文摘要
Large audio-language models (LALMs) demonstrate growing music-understanding capabilities, but whether their responses are grounded in acoustic evidence remains unclear. Musical language often involves abstract concepts whose acoustic evidence is difficult to define and evaluate precisely. We introduce MusicGroundingBench, a controlled benchmark of algorithmically generated piano audio with exact symbolic alignment, comprising three-note and two-bar settings. We evaluate two complementary capabilities: grounding, which localizes the acoustic evidence for a musical query, and understanding, which answers questions about the same excerpts. Our experiments show that cross-modal fine-tuning enables models to learn each capability, but adding grounding supervision does not consistently improve understanding across backbones. We further test whether understanding requires listening through audio-ablation controls that remove or replace the input audio, and use attention analysis to examine whether grounding supervision shifts attention toward note boundaries. Meanwhile, the two evaluated LALMs show limited zero-shot grounding even for basic musical concepts, highlighting grounded music understanding as an important open challenge.