PBSCR:钢琴盗版乐谱作曲家识别数据集
PBSCR: The Piano Bootleg Score Composer Recognition Dataset
- University of Washington(华盛顿大学)
- Columbia University(哥伦比亚大学)
- Harvey Mudd College(哈维穆德学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发布PBSCR数据集,利用IMSLP乐谱与bootleg score二值图像表示,提供9类、100类标注集及未标注预训练集,并给出监督与少样本基线以支持古典钢琴作曲家识别研究。
AI中文摘要:
本文阐述、描述并发布了用于研究古典钢琴音乐作曲家识别的PBSCR数据集。我们的目标是设计一个有助于开展大规模作曲家识别研究的数据集,使其适用于现代架构和训练实践。为实现这一目标,我们利用IMSLP上丰富的乐谱图像和详尽元数据,采用此前提出的名为bootleg score(盗版乐谱)的特征表示来编码符头相对于五线谱线的位置,并以极其简单的格式(二维二值图像)呈现数据,以鼓励快速探索和迭代。该数据集本身包含用于9类识别任务的40,000张62×64的bootleg score图像、用于100类识别任务的100,000张62×64的bootleg score图像,以及用于预训练的29,310张未标注的可变长度bootleg score图像。标注数据以模仿MNIST图像的形式提供,以便极其轻松地进行可视化、处理并高效训练模型。我们纳入了相关信息,将每张bootleg score图像与其底层原始乐谱图像关联起来;同时从IMSLP抓取、整理并汇编了所有钢琴作品的元数据,以促进多模态研究并便于与其他数据集链接。我们发布了监督设置和少样本设置下的基线结果,供未来工作进行比较,并讨论了PBSCR数据特别适合推动研究的开放性研究问题。
英文摘要:
This article motivates, describes, and presents the PBSCR dataset for studying composer recognition of classical piano music. Our goal was to design a dataset that facilitates large-scale research on composer recognition that is suitable for modern architectures and training practices. To achieve this goal, we utilize the abundance of sheet music images and rich metadata on IMSLP, use a previously proposed feature representation called a bootleg score to encode the location of noteheads relative to staff lines, and present the data in an extremely simple format (2D binary images) to encourage rapid exploration and iteration. The dataset itself contains 40,000 62x64 bootleg score images for a 9-class recognition task, 100,000 62x64 bootleg score images for a 100-class recognition task, and 29,310 unlabeled variable-length bootleg score images for pretraining. The labeled data is presented in a form that mirrors MNIST images, in order to make it extremely easy to visualize, manipulate, and train models in an efficient manner. We include relevant information to connect each bootleg score image with its underlying raw sheet music image, and we scrape, organize, and compile metadata from IMSLP on all piano works to facilitate multimodal research and allow for convenient linking to other datasets. We release baseline results in a supervised and low-shot setting for future works to compare against, and we discuss open research questions that the PBSCR data is especially well suited to facilitate research on.