arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19919cs.SD

曲目与版本的统一音乐识别

Unified Music Identification for Tracks and Versions

R. Oguz Araz, Joan Serrà, Yuki Mitsufuji, Xavier Serra, Dmitry Bogdanov

AI总结:

该研究针对曲目与版本识别分开处理的问题,提出统一基准评估7种模型的准确性与鲁棒性,训练出10秒TI查询的基线模型,验证了音乐识别可统一的可行性。

AI中文摘要:

给定一个音乐数据库,曲目识别(TI)用于检索与音频片段匹配的精确曲目,而版本识别(VI)用于检索该曲目的不同音乐版本。传统上,这两项任务是分开处理的。然而,由于每首曲目本身就是其最接近的版本,我们研究是否可以用VI涵盖TI,这要求VI系统对信号操作和音频降质都具有鲁棒性。因此,我们提出了一个统一基准,用于评估每项任务的准确性和鲁棒性。在该基准上对比7种现有模型后,我们发现没有一种模型在两项任务上同时具备准确性和鲁棒性。随后,我们训练了一个针对两项任务的基线模型,结果表明,使用10秒的TI查询,统一系统是可行的。最后,我们分析了限制该模型TI性能的两个检索约束。我们设想将这种统一扩展到其他音乐识别任务。

英文摘要:

Given a music database, track identification (TI) retrieves the exact track matching an audio excerpt, whereas version identification (VI) retrieves its musical versions. Traditionally, the two tasks have been addressed separately. However, as every track is its own closest version, we investigate whether VI can subsume TI. This requires VI systems to be robust to both signal manipulation and audio degradation. We therefore propose a unified benchmark that evaluates accuracy and robustness on each task. Comparing seven existing models on this benchmark, we show that none of them are both accurate and robust on both tasks. We then train a baseline model targeting both tasks and show that a unified system is possible with 10 s TI queries. Lastly, we characterize the two retrieval constraints that limit our model's TI performance. We envision extending this unification to other music identification tasks.

补充信息

↑