arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05816cs.MM

声音从何而来?通过选择性收敛实现自监督双源视听定位

Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

Han Hu, Dongheng Lin, Yuqi Hou, Haotian Li, Hyung Jin Chang, Jianbo Jiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对双源视听定位的循环依赖问题,利用自监督学习的选择性收敛提出两阶段框架,在双源基准上取得自监督方法最优性能,还通过引入像素级掩码修正了基准评估的不一致性。

中文摘要 AI 辅助

在视觉场景中对多个声源进行定位,仍然是多模态感知领域的一项基础挑战,这源于一种固有的循环依赖关系:分离混合音频需要知晓声源位置,而识别发声区域则需要分离后的音频信号。本文聚焦于双源场景,在自监督视听学习中发现了一种选择性收敛现象:当呈现多个声源时,对比模型会自然收敛到最显著的视听对应关系,而非试图平等表征所有声源。这种类似人类选择性听觉注意力的涌现现象,使我们能够通过一个渐进式两阶段框架打破上述循环依赖:首先利用选择性收敛识别主导声源,再利用这些学习到的先验信息挖掘剩余声源。我们的自监督方法在双源基准测试中取得了自监督方法中的最佳性能,无需任何人工标注,在某些指标上甚至超过了一些弱监督方法。此外,我们还发现现有基准测试中存在一个根本性的评估不一致问题:将连续定位热图与边界框标注进行比较会产生系统性偏差,尤其是对于非轴对齐对象,其边界框包含大量背景区域。为解决这一问题,我们在现有基准测试中引入了像素级分割掩码,实现了空间对齐的评估。综上,这些结果表明,接纳而非抑制选择性,为多源定位提供了一条可扩展、无需标注的途径。

英文摘要

Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires knowing source locations, while identifying sound-producing regions requires separated audio signals. In this paper, we focus on the dual-source setting and discover a selective convergence in self-supervised audio-visual learning: when presented with multiple sound sources, contrastive models naturally converge to the most salient audio-visual correspondence rather than attempting to represent all sources equally. This emergent phenomenon, analogous to human selective auditory attention, enables us to break the above circular dependency through a progressive two-stage framework: first, leveraging selective convergence to identify dominant sources, and then exploiting these learned priors to uncover remaining sources. Our self-supervised approach achieves the best performance among self-supervised methods on dual-source benchmarks without requiring any manual annotations, and even surpasses some weakly-supervised approaches \red{on certain metrics. Furthermore, we identify a fundamental evaluation inconsistency in existing benchmarks: comparing continuous localisation heatmaps against bounding-box annotations creates systematic biases, particularly for non-axis-aligned objects where the bounding box includes substantial background regions. To address this, we introduce pixel-level segmentation masks to the existing benchmark, enabling spatially-aligned evaluation. Together, these results suggest that embracing rather than suppressing selectivity offers a scalable, annotation-free route to multi-source localisation.

↑