AI 中文总结
该研究针对真实场景下音乐版本识别的领域不匹配问题,推出大规模DiVers数据集并开展基准测试,训练出的模型在多样含噪输入上鲁棒性显著提升且在干净基准上性能稳定,相关资源已开源支持可复现。
AI 中文摘要
现有的音乐版本识别(VI)数据集主要来自SecondHandSongs和Discogs等经过整理的元数据来源,因此以专业录制的曲目为主,这导致其与业余内容和用户生成内容占主流的真实场景存在领域不匹配问题。为解决这一局限,我们推出DiVers,这是一个包含超过110万个音乐版本的大规模VI数据集,其训练-验证-测试划分与Discogs-VI-YT、SHS100K和Da-TACOS等已建立的数据集兼容。除标准的版本级注释外,DiVers还提供自动分配的标签(如器乐、现场)和指示音乐存在与否的片段级预测。我们通过训练最先进的VI系统来评估所提出的数据集,结果显示,在DiVers上训练的模型对声学多样且含噪声的输入实现了显著提升的鲁棒性,同时在更干净的录音室质量基准上保持稳定性能。我们发布了数据集元数据、其构建代码以及所有实验流程,以支持可复现性。
英文摘要
Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by professionally recorded tracks. This leads to a domain mismatch with real-world scenarios, where amateur and user-generated content is prevalent. To address this limitation, we introduce DiVers, a large-scale VI dataset comprising over 1.1 million musical versions, with train-validation-test splits compatible with established datasets such as Discogs-VI-YT, SHS100K, and Da-TACOS. In addition to standard version-level annotations, DiVers provides automatically assigned tags (e.g., instrumental, live) and segment-level predictions indicating the presence or absence of music. We evaluate the proposed dataset by training state-of-the-art VI systems. Our results show that models trained on DiVers achieve substantially improved robustness to acoustically diverse and noisy inputs, while maintaining a stable performance on cleaner, studio-quality benchmarks. We release the dataset metadata, code for its construction, and all experimental pipelines to support reproducibility.
CommentsAccepted to the Proceedings of the 27th International Society for Music Information Retrieval Conference (ISMIR 2026)