AI 中文总结
该研究推出首个多声部光学音乐识别专用数据集OSSQ-OMR,附带基准协议,测试显示不同编码、分割选择及模型类型对OMR性能影响显著。
AI 中文摘要
光学音乐识别(Optical Music Recognition,OMR)将乐谱转录为数字格式。尽管该领域在单声部和钢琴格式乐谱上已取得显著进展,但多声部乐谱转录仍未得到充分探索,这在很大程度上是由于缺乏合适的数据集。我们推出用于光学音乐识别的OpenScore弦乐四重奏数据集(OpenScore String Quartet for Optical Music Recognition,OSSQ-OMR),这是首个专门针对多声部OMR的数据集。该数据集基于OpenScore弦乐四重奏语料库构建,将数字编码乐谱与来自IMSLP的原始扫描版本配对,所有图像均与其转录内容视觉对齐。数据集发布了系统级和五线谱级的乐谱图像,以及三种编码格式的配对转录内容:扩展线性化MusicXML(Extended Linearized MusicXML,LMXE)、**kern和ABC。总体而言,OSSQ-OMR包含来自116份弦乐四重奏乐谱的24544张系统图像和98172张五线谱图像。我们随数据集提供基准测试协议,以及两种代表性OMR模型的基线结果,这些结果在四个具有互斥测试集的随机乐谱级划分上进行评估。基线在合成输入上达到OMR-NED低至3.6%,在扫描输入上达到5.9%;结果显示编码和分割选择产生显著影响,基于LSTM的基线在扫描输入上的性能下降幅度约为基于Transformer的基线的2.6倍。
英文摘要
Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form scores, multi-part score transcription remains underexplored, largely due to the absence of a suitable dataset. We introduce OpenScore String Quartet for Optical Music Recognition (OSSQ-OMR), the first dataset dedicated to multi-part OMR. Built on the OpenScore String Quartet corpus, OSSQ-OMR pairs digitally encoded scores with their original scanned editions from IMSLP, with all images visually aligned to their transcriptions. The dataset is released with score images at system and staff levels, and paired transcriptions in three encoding formats: Extended Linearized MusicXML (LMXE), **kern, and ABC. In total, OSSQ-OMR contains 24,544 system images and 98,172 staff images drawn from 116 string quartet scores. We accompany the dataset with a benchmark protocol and baseline results from two representative OMR models, evaluated across four random score-level splits with mutually exclusive test sets. Baselines reach OMR-NED as low as 3.6% on synthetic and 5.9% on scanned inputs; results reveal substantial effects of encoding and segmentation choices, with the LSTM-based baseline degrading on scanned inputs roughly 2.6 times less than the Transformer-based baseline.
Comments8 pages, 2 figures, 5 tables. Accepted at the ISMIR 2026