arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15465cs.SD

构建音乐样本识别数据集

Building a Dataset for Music Sample Identification

R. Oguz Araz, Xavier Lizarraga, Xavier Serra, Dmitry Bogdanov

首次发表
浏览论文内容

中文总结 AI 辅助

针对样本识别任务缺乏大规模数据的问题,构建了一个比现有基准大近三个数量级的数据集,并提出了基于图连通分量划分的防泄漏流水线,以促进该领域研究。

中文摘要 AI 辅助

样本识别(SI)是将音乐作品中的元素与其被音乐化变换后用于创作新作品的版本进行匹配的任务。该任务受到的关注较少,且缺乏大规模公开可用的数据。在本工作中,我们从音乐数据库中挖掘采样注释,并将其划分为训练集和评估集。由此产生的数据集比现有的SI基准大了近三个数量级,其中训练集、验证集和测试集分别包含114,000、6,000和10,000首曲目。我们发现,简单地对注释进行划分会将相同的曲目放入不同的集合中。为避免这种情况,我们根据注释构建了一个图,并在其连通分量上进行划分。我们进一步发现,一个单一的巨型分量包含了注释的一半,这使得按分量划分与平衡划分不兼容;我们对其进行了修剪,从而产生了一个防泄漏的流水线。我们仅出于非商业科学研究目的共享该数据集,并公开数据分析和划分代码。我们希望我们的工作能促进对SI的研究。

英文摘要

Sample identification (SI) is the task of matching an element of a musical work to its musically transformed versions used to create new works. The task has received little attention and lacks large-scale publicly available data. In this work, we mine sampling annotations from a music database and split them for training and evaluation. The resulting dataset is nearly three orders of magnitude larger than the existing SI benchmarks, with training, validation, and test sets of 114 k, 6 k, and 10 k tracks. We find that naively splitting the annotations places the same tracks in different sets. To avoid this, we construct a graph from the annotations and split it over connected components. We further find that a single mega-component contains half of the annotations, making component-wise splitting incompatible with balanced splits; we trim it, yielding a leakage-aware pipeline. We share the dataset for non-commercial scientific research purposes only and make the data-analysis and splitting code publicly available. We hope that our work fosters research on SI.

补充信息

↑