arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05662cs.CVcs.AI

手写单声部乐谱的全页光学音乐识别

Full-Page Optical Music Recognition of Handwritten Monophonic Scores

  • Pattern Recognition and Artificial Intelligence group, University of Alicante(阿利坎特大学模式识别与人工智能组)
  • Instituto Superior de Enseñanzas Artísticas de la Comunidad Valenciana(瓦伦西亚自治区高等艺术教育机构)

机构由 AI 辅助整理,请以论文原文为准。

Adrian Rosello, Antonio Ríos-Vila, David Rizo, Jorge Calvo-Zaragoza

AI总结:

本研究探索手写单声部乐谱的全页端到端光学音乐识别,通过引入合成数据生成器分析预训练因素,实验表明合成预训练主要受益于结构布局学习而非视觉相似性。

AI中文摘要:

全页端到端光学音乐识别旨在将整个音乐页面直接转录为符号表示,从而避免依赖精确谱表分割的传统流水线的局限性。近年来,基于Transformer的架构在排版乐谱上取得了强劲性能,其预训练依赖于大规模合成数据。然而,这些方法对手写音乐的适用性在很大程度上尚未被探索。在本工作中,我们研究了手写单声部乐谱集合上的全页转录,并分析了合成预训练在此场景下的影响。为了探究预训练期间哪些因素最为关键,我们引入了一个能够生成视觉连贯的全页乐谱的生成器,该生成器支持排版风格和手写风格两种形式。在三个真实手写数据集上的实验提供了对多种全页流水线和不同合成预训练策略的比较评估。结果表明,合成预训练带来的益处主要与学习结构布局惯例相关,而非与目标手写体的视觉相似性相关。

英文摘要:

Full-page end-to-end Optical Music Recognition seeks to transcribe entire music pages directly into symbolic notation, avoiding the limitations of traditional pipelines that rely on accurate staff segmentation. Recent Transformer-based architectures have achieved strong performance on typeset scores, relying on large-scale synthetic data for pretraining. However, their applicability to handwritten music remains largely unexplored. In this work, we study full-page transcription on handwritten monophonic collections and analyze the impact of synthetic pretraining in this setting. To investigate which factors are most relevant during pretraining, we introduce a generator capable of producing visually coherent full-page scores in both typeset and handwritten styles. Experiments on three real handwritten datasets provide a comparative evaluation of several full-page pipelines and different synthetic pretraining strategies. The results suggest that the benefits of synthetic pretraining are primarily associated with learning structural layout conventions rather than with visual similarity to the target handwriting.

↑