发表机构
Multimodal Art Projection; New York University; MBZUAI; ACE Studio; The Hong Kong University of Science and Technology(多模态艺术投影; 纽约大学; 穆罕默德·本·扎耶德人工智能大学; ACE Studio; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对音乐转录中标注稀缺和局部预测不一致的问题,提出统一框架SheetSage2,结合合成数据、结构化解码和自回归蒸馏,在八个基准上超越多个先前系统,实现连贯的主旋律谱转录。
AI 中文摘要
将音乐转录为人类可读的乐谱需要对节奏、和声、旋律和曲式有连贯的理解。两个障碍限制了这一目标:带标注的录音稀缺,且准确的局部预测仍可能产生不一致的音乐序列。我们提出了SheetSage2,一个统一的音乐转录框架,它结合了合成数据、任务特定的结构化解码和自回归蒸馏。自动标注的MIDI被渲染为音频,为音乐理解任务提供了可扩展的监督。任务特定的结构化解码器整合互补的音乐线索及其时间依赖关系,以生成音乐上连贯的乐谱。自回归蒸馏进一步在推理时无需任务特定的动态规划即可保持转录准确性。在八个基准集合中,单个SheetSage2-AR模型在我们的评估中超过了15个基准-指标对中的12个所列先前系统,显著优于SheetSage1,并在多个基准上超越了任务特定模型。模型权重和推理代码已公开提供。
英文摘要
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.