arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TUTTI:基于完全合成数据实现可泛化的音频到乐谱转录

TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun

arXiv 2609.00640首次发表:更新:

AI 中文总结

TUTTI是一种基于合成多乐器数据训练的A2S预训练范式,可提升模型泛化能力,在多乐器A2S任务上达新SOTA,且跨乐器迁移性优异。

AI 中文摘要

可泛化的音频到乐谱(Audio-to-Score,A2S)转录本质上受限于高质量真实配对数据的严重稀缺。仅依赖现有人工标注数据集通常会限制A2S模型的泛化能力,使其效能主要局限于单乐器领域。为打破对稀缺真实数据的依赖,我们提出TUTTI(Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data,基于合成多乐器数据训练的统一音频到乐谱转录Transformer),这是一种由纯合成大规模数据集驱动的预训练范式。我们未使用人工创作的乐谱,而是利用符号音乐生成模型生成海量、高度可扩展的多乐器语料库,并创建具有丰富声学特征的音频-乐谱配对数据。基于生成的数据,我们采用标准Transformer编码器-解码器架构。我们通过实验证明,在生成的多乐器数据上预训练基于注意力的统一模型,相比单乐器训练能获得更一致的强基础表示。当用下游真实世界数据集微调时,TUTTI优于现有方法,在各类A2S基准中建立了新的整体最优结果。值得注意的是,TUTTI表现出显著的跨乐器迁移能力,能有效适配未见乐器并取得极具竞争力的性能。源代码和TuttiCorpus数据集将在该公开URL上发布。

英文摘要

Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiting their efficacy primarily to single-instrumentation domains. To break this dependency on scarce real-world data, we introduce TUTTI (Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data), a pre-training paradigm driven by a purely synthetic, large-scale dataset. Rather than using human-composed scores, we leverage a symbolic music generation model to generate a massive, highly scalable multi-instrumentation corpus and create audio-score pairs with expressive acoustic characteristics. Capitalizing on the generated data, we employ a standard Transformer encoder-decoder architecture. We empirically demonstrate that pre-training a unified attention-based model on generated, multi-instrumentation data yields a consistently stronger foundational representation than single-instrumentation training. When fine-tuned with downstream real-world datasets, TUTTI outperforms previous approaches, establishing new overall state-of-the-art results across various A2S baselines. Notably, TUTTI shows remarkable cross-instrument transferability, effectively adapting to unseen instruments with highly competitive performance. The source code and the TuttiCorpus dataset will be made publicly available at https://github.com/a-musiclover/TUTTI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑