AI 中文总结
TUTTI是一种基于合成多乐器数据训练的A2S预训练范式,可提升模型泛化能力,在多乐器A2S任务上达新SOTA,且跨乐器迁移性优异。
AI 中文摘要
可泛化的音频到乐谱(Audio-to-Score,A2S)转录本质上受限于高质量真实配对数据的严重稀缺。仅依赖现有人工标注数据集通常会限制A2S模型的泛化能力,使其效能主要局限于单乐器领域。为打破对稀缺真实数据的依赖,我们提出TUTTI(Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data,基于合成多乐器数据训练的统一音频到乐谱转录Transformer),这是一种由纯合成大规模数据集驱动的预训练范式。我们未使用人工创作的乐谱,而是利用符号音乐生成模型生成海量、高度可扩展的多乐器语料库,并创建具有丰富声学特征的音频-乐谱配对数据。基于生成的数据,我们采用标准Transformer编码器-解码器架构。我们通过实验证明,在生成的多乐器数据上预训练基于注意力的统一模型,相比单乐器训练能获得更一致的强基础表示。当用下游真实世界数据集微调时,TUTTI优于现有方法,在各类A2S基准中建立了新的整体最优结果。值得注意的是,TUTTI表现出显著的跨乐器迁移能力,能有效适配未见乐器并取得极具竞争力的性能。源代码和TuttiCorpus数据集将在该公开URL上发布。
英文摘要
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiting their efficacy primarily to single-instrumentation domains. To break this dependency on scarce real-world data, we introduce TUTTI (Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data), a pre-training paradigm driven by a purely synthetic, large-scale dataset. Rather than using human-composed scores, we leverage a symbolic music generation model to generate a massive, highly scalable multi-instrumentation corpus and create audio-score pairs with expressive acoustic characteristics. Capitalizing on the generated data, we employ a standard Transformer encoder-decoder architecture. We empirically demonstrate that pre-training a unified attention-based model on generated, multi-instrumentation data yields a consistently stronger foundational representation than single-instrumentation training. When fine-tuned with downstream real-world datasets, TUTTI outperforms previous approaches, establishing new overall state-of-the-art results across various A2S baselines. Notably, TUTTI shows remarkable cross-instrument transferability, effectively adapting to unseen instruments with highly competitive performance. The source code and the TuttiCorpus dataset will be made publicly available at https://github.com/a-musiclover/TUTTI.