发表机构
University of Washington; Allen Institute for Artificial Intelligence(华盛顿大学; 艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对音乐转录系统提出记谱相似度与回放相似度的双重评估框架,发现CLEWS指标相关性佳且成本低,不同评估维度偏好不同系统,新系统Rubato记谱相似度提升且回放具竞争力。
AI 中文摘要
自动音乐转录系统生成可阅读和回放的乐谱。我们认为这两个目标需要分别对与参考乐谱的记谱相似度、与原始演奏的回放相似度进行互补评估。本研究考虑来自光学音乐识别文献的记谱相似度指标,以及通过针对100余名参与者、230份钢琴录音(涵盖23部作品、30位演奏者和6位作曲家)的聆听研究验证的多种回放相似度方法。我们意外发现,与人类判断相关性最佳的回放相似度指标CLEWS,同时也是运行成本最低的指标。我们还发现,在由8种音频转MIDI模型与3种MIDI转乐谱转换器配对形成的24条流水线集合中,两个评估维度偏好不同的系统,且后者组件系统性地决定了偏好的目标。当加入全新的端到端系统Rubato时,指标间的互补性依然成立:该系统的记谱相似度显著提升,同时回放相似度仍具有竞争力(虽非最优)。
英文摘要
Automatic music transcription systems produce sheet music that can be read and played back. We argue that these two targets call for complementary evaluations of notation similarity to a reference score and playback similarity to the original performance, respectively. Our study considers notation similarity metrics from the optical music recognition literature and a wide range of playback-similarity methods validated through a listening study across over 100 participants and 230 piano recordings covering 23 works, 30 performers, and six composers. We find, fortuitously, that the playback similarity metric that correlates best with human judgments, CLEWS, is also the cheapest to run. We also find that the two evaluation dimensions favor different systems among a collection of 24 pipelines formed by pairing eight audio-to-MIDI models with three MIDI-to-score converters, with the latter component systematically determining the favored objective. The complementarity between metrics also holds when adding to the pool Rubato, a new end-to-end system that offers substantially improved notation similarity while remaining competitive, though not the best, on playback similarity.