arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在大型视听检索模型中解锁空间定位

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

Hugo Malard, Michel Olvera, Sanjeel Parekh, Gaël Richard, Slim Essid, Stéphane Lathuilière

arXiv 2607.24786首次发表:更新:

发表机构

LTCI, Télécom Paris, Institut Polytechnique de Paris; Meta, Reality Labs Research; NVIDIA; Inria at Université Grenoble Alpes, CNRS, LJK(巴黎综合理工学院电信与信号处理实验室,巴黎电信学院; 元宇宙公司 Reality Labs 研究部; 英伟达公司; 法国国家信息与自动化研究所,格勒诺布尔阿尔卑斯大学,法国国家科学研究中心,让·昆汀·莫诺德实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视听声源定位任务,利用大规模视听检索模型的潜在表示,引入LAIP框架,通过音频信息池化恢复局部空间信息,在相关数据集上取得领先性能,证明可从现有检索表示解锁定位,为检索和定位任务提供统一路径。

AI 中文摘要

弱监督为视听声源定位设定了一种实用模式,因为大规模获取密集的空间注释成本很高。然而,该任务仍然具有挑战性,因为模型必须在没有像素级监督的情况下从时间对齐的视听数据中定位声源。最近训练的大规模视听检索模型编码了丰富的多模态结构。我们表明,其潜在表示虽然针对全局对齐进行了优化,但仍能实现细粒度的空间定位。由于全局池化,空间细节在检索主干的上层逐渐丢失,但中间视觉令牌保留了高度结构化的空间信息。为利用此信息,我们引入了LAIP(通过音频信息池化进行定位)框架,该框架采用轻量级的音频信息空间池化(AiSP)来替换标准的全局聚合模块。通过使用帧对齐音频查询中间视觉令牌,LAIP恢复了否则会被冻结的检索管道丢弃的局部空间信息。我们的方法在AVSBench和AVATAR上取得了领先的性能,在后者上的结果几乎是之前的两倍。这些发现证明,准确的定位不需要从头开始学习;相反,可以从现有的检索表示中解锁,为检索和定位任务提供了统一的路径。

英文摘要

Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By querying intermediate visual tokens with audio aligned at the frame level, LAIP recovers localized spatial information that is otherwise discarded by the retrieval pipeline, with the largest gains observed for PE-AV, a stack with underlying temporal aggregation. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter, improving average CIoU from 13.21 to 26.22.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑