CODA:用于基于图像的实时乐谱跟随的级联在线不连续感知对齐
CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following
浏览论文内容
中文总结 AI 辅助
本文针对实时乐谱跟随难题,提出CODA系统。它利用乐谱级联结构增强预测一致性,通过静音驱动中断模式实现不连续恢复。在多模态乐谱数据集钢琴基准测试中,CODA在实时吞吐量下取得了领先的跟踪精度和不连续恢复性能。
中文摘要 AI 辅助
从乐谱图像进行实时乐谱跟随具有挑战性,因为模型必须在严格延迟约束下处理流音频并解决高度重复的视觉模式。近期基于图像的方法尝试通过同时预测活动系统、小节和音符的位置来使用多分辨率预测,但不同符号级别的预测相互独立,导致预测不稳定并引入不必要的额外搜索空间,且多数现有方法缺乏从乐谱不连续(如重复、从头反复或尾声跳转)中恢复的机制。本文提出CODA,这是首个解决上述两个问题的实时乐谱跟随系统。CODA明确利用乐谱的级联结构,先选择活动系统,再选其中的活动小节,最后选所选小节内的活动音符,从而增强跨分辨率的预测一致性。一种由静音驱动的中断模式可从任意乐谱不连续中恢复,而无需重复结构的知识。在多模态乐谱数据集(MSMD)钢琴基准测试中评估,CODA在实时吞吐量下实现了领先的跟踪精度和不连续恢复性能。代码可在指定网址获取。
英文摘要
Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction by simultaneously predicting the positions of the active sys- tem, bar, and note. However, their predictions across these different levels of notation are independent, which makes the predictions unstable and introduces unnecessary ex- tra search space for bar- and note-level predictions. Most existing methods also lack mechanisms to recover from score discontinuities, such as repeats, da capo (D.C.), or coda jumps. This paper proposes CODA, to the best of our knowledge, the first real-time score following system that addresses both gaps. CODA explicitly exploits the cascaded structure of music scores: it first selects the ac- tive system, then the active bar within it, and finally the active note within the selected bar. This enforces pre- diction consistency across resolutions. A silence-driven break mode enables recovery from arbitrary score discon- tinuities without requiring knowledge of the repeat struc- ture. Evaluated on the Multimodal Sheet Music Dataset (MSMD) piano benchmarks, CODA achieves state-of-the- art tracking accuracy and discontinuity-recovery perfor- mance under real-time throughput. Code is available at https://github.com/ValleyC/CODA.