arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于预训练特征、数据增强及新SheetSage-A2S数据集的音频转乐谱转录

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoru Mo, Yaolong Ju

arXiv 2608.06165首次发表:更新:

发表机构

University College Dublin; Guangxi Normal University; Great Bay University; Shenzhen University(都柏林大学学院; 广西师范大学; 大湾区大学; 深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对流行音乐音频转乐谱研究不足的问题,构建了SheetSage-A2S数据集,结合预训练模型MuQ与数据增强改进A2S方法,在古典与流行音乐基准上均取得优于现有技术的性能。

AI 中文摘要

现有的音频转乐谱(A2S)系统主要聚焦于古典音乐,在流行音乐中的应用仍未得到充分探索。本文首先推出新的SheetSage-A2S数据集,该数据集包含6066首独特歌曲的9468个片段,共61小时音频及对应的\texttt{**kern}乐谱编码,是首个助力流行音乐A2S研究的同类数据集。此外,我们通过数据增强和音乐音频预训练特征提取模型MuQ改进现有A2S方法,以提升泛化能力并提取有意义的音频特征。结果显示,所提A2S模型在古典音乐的Quartets数据集上实现4.98%的符号错误率(SER),显著优于现有最先进方法的15.3% SER;在流行音乐的SheetSage-A2S数据集上实现20.92%的SER,为未来研究提供了强基准。该数据集、模型及代码可在此公开获取。

英文摘要

Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with **kern score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3% SER from the existing state-of-the-art (Alfaro-Contreras et al. 2024). Additionally, our model achieves 20.92% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern_model.

CommentsAccepted at the 34th ACM International Conference on Multimedia (MM '26) 2026-08-25: Edit to drop TeX commands in abstract

DOI:10.1145/3767308.3835653

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑