arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

POLARIS:基于显著性地标与Delaunay分组的免训练音频指纹识别

POLARIS: Training-Free Audio Fingerprinting with Saliency-Based Landmarks and Delaunay Grouping

Jiheng Li

arXiv 2609.14820首次发表:更新:

发表机构

Vanderbilt University(范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

POLARIS是一种免训练的音频指纹系统,利用显著性地标和Delaunay分组生成稀疏指纹,通过自适应两跳邻域扩展应对失真,在合成和真实重录基准上优于现有免训练方法及神经基线。

AI 中文摘要

本文提出POLARIS,一种免训练的音频指纹识别系统,该系统从局部归一化显著性场中选择地标,并利用Delaunay三角剖分将其分组为稀疏指纹。为应对查询失真,POLARIS仅在查询时添加来自两跳Delaunay邻域的指纹,而不扩大参考索引。一种自适应配置仅在原始指纹未能产生可信匹配时应用此扩展。我们在公开的PEX Hard Medium基准的合成失真(排除音高或速度偏移的查询)以及一个新的真实重录音乐基准上评估POLARIS。POLARIS在两个基准上均达到所评估的免训练方法中的最佳性能。在真实录音上,其自适应配置也优于神经NMFP基线,同时查询时间相当且逻辑参考负载更小。代码、数据集及复现所有实验的说明可在该https URL获取。

英文摘要

This work presents POLARIS, a training-free audio fingerprinting system that selects landmarks from a locally normalized saliency field and groups them into sparse fingerprints using Delaunay triangulation. To deal with query distortion, POLARIS adds fingerprints from two-hop Delaunay neighborhoods only at query time, without enlarging the reference index. An adaptive configuration applies this expansion only when the original fingerprints do not produce a confident match. We evaluate POLARIS on synthetic distortions from the public PEX Hard Medium benchmark, excluding queries with pitch or tempo shifts, and on a new benchmark of real re-recorded music. POLARIS achieves the best performance among the evaluated training-free methods on both benchmarks. On the real recordings, its adaptive configuration also outperforms the neural NMFP baseline with a comparable measured query time and a smaller logical reference payload. Code, dataset, and instructions for reproducing all experiments are available at https://github.com/JihengLi/POLARIS.git.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑