音频能否揭示乐曲演奏难度?来自Piano Syllabus数据集的见解
Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset
- Music Technology Group, Universitat Pompeu Fabra(庞培法华大学音乐技术小组)
- Music & Art Learning Lab, Sogang University(ソガン大学音乐与艺术学习实验室)
- Pattern Recognition and Artificial Intelligence Group, University of Alicante(阿尔基兰特大学模式识别与人工智能小组)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究首创基于音频录音的乐曲演奏难度自动评估方法,构建了首个包含7,901首作品的Piano Syllabus数据集,并提出支持单模态与多模态输入的识别框架,实验证明其有效性。
AI中文摘要:
自动评估乐曲的演奏难度是音乐教育中根据学生个人需求制定量身定制课程的关键过程。鉴于其重要性,音乐信息检索(MIR)领域展示了一些针对该任务的概念验证工作,这些工作主要关注机器可读乐谱或乐谱图像等高级音乐抽象表示。在这方面,直接分析音频录音的潜力通常被忽视,这阻碍了学生探索可能没有正式符号层面转录的多样化乐曲。本研究开创性地在音频录音上自动评估乐曲演奏难度,并做出了两项具体贡献:(i) 首个基于音频的难度评估数据集——即 Piano Syllabus (PSyllabus) 数据集——包含来自 1,233 位作曲家的 7,901 首钢琴作品,涵盖 11 个难度级别;以及 (ii) 一个能够管理不同输入表示(包括单模态和多模态方式)的识别框架,这些表示直接源自音频以执行难度评估任务。包含不同预训练方案、输入模态和多任务场景的综合实验证明了该提议的有效性,并确立了 PSyllabus 作为 MIR 领域中基于音频的难度评估的参考数据集。数据集以及开发的代码和训练好的模型均已公开分享,以促进该领域的进一步研究。
英文摘要:
Automatically estimating the performance difficulty of a music piece represents a key process in music education to create tailored curricula according to the individual needs of the students. Given its relevance, the Music Information Retrieval (MIR) field depicts some proof-of-concept works addressing this task that mainly focuses on high-level music abstractions such as machine-readable scores or music sheet images. In this regard, the potential of directly analyzing audio recordings has been generally neglected, which prevents students from exploring diverse music pieces that may not have a formal symbolic-level transcription. This work pioneers in the automatic estimation of performance difficulty of music pieces on audio recordings with two precise contributions: (i) the first audio-based difficulty estimation dataset -- namely, Piano Syllabus (PSyllabus) dataset -- featuring 7,901 piano pieces across 11 difficulty levels from 1,233 composers; and (ii) a recognition framework capable of managing different input representations -- both unimodal and multimodal manners -- directly derived from audio to perform the difficulty estimation task. The comprehensive experimentation comprising different pre-training schemes, input modalities, and multi-task scenarios prove the validity of the proposal and establishes PSyllabus as a reference dataset for audio-based difficulty estimation in the MIR field. The dataset as well as the developed code and trained models are publicly shared to promote further research in the field.