arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28212cs.MMcs.SD

MuSP-Bench:面向乐谱与演奏的高级多模态音乐理解基准测试

MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance

Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员推出MuSP-Bench基准测试,含490个乐谱与演奏理解问题,评估发现前沿多模态大语言模型理解乐谱困难,推理演奏音频时挑战更大。

中文摘要 AI 辅助

音乐家通常通过乐谱和演奏来交流音乐,乐谱编码音乐意图,演奏则将其以声音形式实现。为探究模型能否有效处理这两种模态,我们推出MuSP-Bench,这是一个由人工编写的基准测试,包含490个针对音乐乐谱与演奏理解的问题。该基准测试的独特之处在于覆盖古典钢琴和管弦乐作品中基于乐谱、基于演奏、诠释性及长时程推理任务。我们在多种输入条件下评估前沿多模态大语言模型,结果显示这些模型在理解乐谱方面存在显著困难,在对演奏音频进行推理时面临更大挑战。该基准测试可在指定网址获取。

英文摘要

Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.

↑