双耳视听实例分割
Binaural Audio-Visual Instance Segmentation
浏览论文内容
中文总结 AI 辅助
本文提出双耳视听实例分割(BiAVIS)任务,利用双耳音频的空间线索和查询级视听融合,解决单耳AVS中同类实例难以区分的问题,并在新基准上显著提升性能。
中文摘要 AI 辅助
视听分割(AVS)旨在通过整合听觉和视觉线索,在像素级别分割出发声物体。然而,现有方法主要是在单耳设置下开发的,并且主要依赖于跨模态语义对应,这限制了它们区分同一语义类别中视觉相似实例的能力。相比之下,人类自然地利用双耳听觉,其中由头部和耳廓引入的耳间差异和方向相关的声学滤波为准确的声源定位提供了基于物理的空间线索。受此观察的启发,我们引入了双耳视听实例分割(BiAVIS),这是一个利用同步双耳音频和视频帧来分割发声实例的新任务。为了推进这一任务的研究,我们通过手动标注现有的双耳视听数据集并收集一个新的真实世界数据集BiAVIS-Bench(在更具挑战性和多样性的场景中)来建立两个基准。我们进一步提出了一种BiAVIS模型,该模型利用纯音频声源定位网络从双耳音频中学习发声实例的空间和语义先验。随后引入了一种查询级别的视听融合策略,将这些信息丰富的先验注入实例分割解码器。在两个提出的基准上进行的广泛实验表明,BiAVIS模型优于之前的单耳AVS方法,特别是在解决实例级别的类内歧义方面。在更具挑战性的BiAVIS-Bench上,所提出的BiAVIS模型在mAP和FSLA上分别比最佳性能的单耳基线高出17.22%和7.71%。
英文摘要
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localization. Motivated by this observation, we introduce binaural audio-visual instance segmentation (BiAVIS), a new task that leverages synchronized binaural audio and video frames to segment sounding instances. To advance research on this task, we establish two benchmarks by manually annotating an existing binaural audio-visual dataset and collecting a new real-world dataset, BiAVIS-Bench, in more challenging and diverse scenarios. We further propose a BiAVIS model, which leverages an audio-only sound source localization network to learn spatial and semantic priors for sounding instances from binaural audio. A query-level audio-visual fusion strategy is subsequently introduced to inject these informative priors into the instance segmentation decoder. Extensive experiments conducted on the two proposed benchmarks demonstrate the superior performance of the BiAVIS model over previous monaural AVS methods, especially in resolving instance-level intra-class ambiguity. On the more challenging BiAVIS-Bench, the proposed BiAVIS model outperforms the best-performing monaural baselines by 17.22\% in mAP and 7.71\% in FSLA, respectively.