arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21911cs.SD

基于真实样本对的音乐样本识别模型训练

Training Music Sample Identification Models on Real Sample Pairs

  • Universitat Pompeu Fabra(庞培法布拉大学)
  • Sony Europe(索尼欧洲)
  • Sony CTC America(索尼CTC美洲)
  • BMAT Licensing S.L.(BMAT许可公司)

机构由 AI 辅助整理,请以论文原文为准。

R. Oguz Araz, Joan Serrà, Xavier Lizarraga-Seijas, Xavier Serra, Yuki Mitsufuji, Dmitry Bogdanov

AI总结:

本文提出SI嵌入模型,利用真实样本对训练,在三个基准上达到最先进性能,并首次提供完全监督训练方法,为样本识别领域奠定基础。

AI中文摘要:

样本识别(SI)是一项匹配音轨对的任务,其中一条音轨是通过对另一条音轨的元素进行音乐变换而创建的。在缺乏大规模样本标注的情况下,主导的训练范式一直依赖于人工创建样本对。尽管最近发布的数据集提供了大规模真实样本对的标注,但仍缺少有效的训练方法。在这项工作中,我们提出了SI嵌入(SIE),一种在三个基准测试(包括一个大规模测试集)上达到最先进结果的SI模型。我们表明,先前在人工样本对上训练的最先进模型仅能部分泛化到真实样本对,并且其训练数据限制了其性能。我们还表明,真实样本对并不能完全解释SIE的性能:其架构和训练方法贡献显著。我们提供了首个用于真实世界SI的完全监督训练方法,为该领域的未来研究奠定了坚实基础。

英文摘要:

Sample identification (SI) is the task of matching pairs of tracks, where one track is created by musically transforming an element of the other. In the absence of sample annotations at scale, the dominant training paradigm has depended on artificially creating sample pairs. Although a recently released dataset provides annotations of real sample pairs at scale, an effective training recipe is missing. In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set. We show that the previous state of the art trained on artificial pairs generalizes only partially to real pairs, and that its training data limits its performance. We also show that real pairs do not fully account for SIE's performance: its architecture and training recipe contribute substantially. We provide the first fully supervised training recipe for real-world SI, establishing a strong foundation for future research in the field.

↑