统一乐谱与演奏:面向音频语言模型的细粒度音乐理解
Unifying Score and Performance for Fine-Grained Music Understanding in Audio-Language Models
浏览论文内容
中文总结 AI 辅助
针对音频语言模型在细粒度音乐理解上的不足,提出统一乐谱与演奏的MuNo-SP表示及自动数据生成流程,构建MAESTROCaps数据集,经人工评估和基准测试验证其优于MIDI基线。
中文摘要 AI 辅助
大型音频语言模型(LALMs)在广泛的音乐理解任务(如标签分类、检索和字幕生成)中已展现出令人瞩目的进展。然而,需要对内容及其在演奏中如何通过力度、乐句、运音、时值及其他演奏技巧得以实现进行更精细聆听的音乐理解,仍处于早期阶段。现有的音频语言模型(ALM)训练流程通常依赖粗略、弱语义关联的字幕,因此难以支持学习音乐中的这些细微差别,限制了其在教育或艺术实践等现实应用中的能力。为此,我们引入了MuNo-SP(音乐记谱统一乐谱与演奏),一种基于文本的表示方法,可联合编码乐谱内容和演奏信息。基于MuNo-SP,我们开发了一条自动训练数据生成流程,利用对齐的乐谱和演奏来生成长篇听觉分析及具有音乐知识的问题-答案对。我们利用该流程构建了MAESTROCaps,一个古典钢琴数据集,包含148篇长篇演奏分析和由148对对齐的乐谱-演奏对导出的31,080个问题-答案对。在人工评估中,对于九个片段中的八个,MuNo-SP分析以多数票优于仅基于MIDI的分析。MuNo-SP在乐谱-演奏理解基准测试中也表现强劲,表明整合乐谱和演奏信息比仅基于MIDI的基线能为LALM提供更可靠且更具音乐信息量的监督。
英文摘要
Large audio language models (LALMs) have shown promising progress in broad music-understanding tasks such as tagging, retrieval, and captioning. Music understanding that requires finer hearing over both the content and how it is realized within a performance through dynamics, phrasing, articulation, time, and other performance techniques, however, remains at an earlier stage. Existing audio-language model (ALM) training pipelines typically rely on coarse, weakly grounded captions and therefore provide little support for learning these subtle nuances in music, limiting their ability to serve real-world applications in education or artistic practice. We therefore introduce MuNo-SP (Music Notation unifying Score and Performance), a text-based representation that jointly encodes score content and performance information. Building on MuNo-SP, we develop an automatic training-data generation pipeline that uses aligned scores and performances to produce long-form auditory analyses and musically informed question-answer pairs. We use this pipeline to construct MAESTROCaps, a classical piano dataset comprising 148 long-form performance analyses and 31,080 question-answer pairs derived from 148 aligned score-performance pairs. In a human evaluation, MuNo-SP analyses were preferred by majority vote over MIDI-only analyses for eight of nine excerpts. MuNo-SP also performed strongly on a benchmark of score-performance understanding, suggesting that integrating score and performance information enables more reliable and musically informative LALM supervision than a MIDI-only baseline.
发表机构
- Bryel Labs(Bryel 实验室)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。