AI 中文总结
研究多乐器音乐转录问题,核心方法是分析合成数据预训练有效性,结合真实音乐音频微调与强化学习后训练,还引入乐器存在条件定制转录,主要贡献是发布MuScriptor开放模型用于多种音乐录音。
AI 中文摘要
现有的自动音乐转录方法通常局限于单乐器录音,或在复杂的真实音乐混音中失败。尽管之前的工作使用了合成训练数据,但生成的模型泛化能力差,在现实的多乐器设置中导致转录输出基本无法使用。在这项工作中,我们分析了合成数据预训练的有效性,同时将其与对真实音乐音频的微调以及强化学习的后训练相结合。我们还引入了基于乐器存在的条件来定制转录。最后,我们发布了MuScriptor,一个开放权重的多乐器音乐转录模型,可用于各种音乐类型的真实世界音乐录音。
英文摘要
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.
CommentsISMIR 2026 Camera Ready