基于深度引导协同建模的视听分割
Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
浏览论文内容
中文总结 AI 辅助
该研究针对现有视听分割方法未建模几何线索的问题,提出三模态框架DGCM-AVS,通过深度感知动态调制器与深度引导渐进融合模块优化,在AVSS数据集上实现M_J、M_F指标的显著提升。
中文摘要 AI 辅助
视听分割(Audio-Visual Segmentation, AVS)是多模态感知领域的基础任务,通过结合视觉与音频线索对视频中发声对象进行像素级分割,在视频理解、人机交互、自动驾驶等领域有广泛应用。然而,现有大多数AVS方法未显式建模相对距离、遮挡等几何线索,限制了跨模态对齐的鲁棒性。人类感知中,空间结构会自然与视听证据结合以精准定位发声对象,受此启发,本文将估计的深度作为空间结构线索融入AVS任务,提出三模态框架DGCM-AVS,联合建模音频、视觉与深度信息。具体而言,本文设计了深度感知动态调制器,以提升相邻对象的分离能力,同时保持对象内部特征一致性;还提出了深度引导渐进融合模块,利用深度作为中间桥梁逐步对齐音频线索与视觉特征。在AVSS数据集上,DGCM-AVS相较于现有最优方法,在M_J指标上取得了10.2%的相对提升,在M_F指标上取得了8.7%的相对提升。本文的研究表明,深度是AVS领域极具潜力但尚未充分探索的模态,有望推动该方向的进一步研究。
英文摘要
Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer interaction, and autonomous driving. However, most existing AVS methods do not explicitly model geometric cues such as relative distance and occlusion, thereby limiting the robustness of cross-modal alignment. In human perception, spatial structure is naturally integrated with audio-visual evidence to accurately localize sounding objects. Motivated by this, we incorporate estimated depth as a spatial structural cue for AVS and propose DGCM-AVS, a tri-modal framework that jointly models audio, visual, and depth information. Specifically, we design a Depth-Aware Dynamic Modulator to improve the separation of adjacent objects while preserving intra-object feature consistency. Furthermore, we propose Depth-Guided Progressive Fusion, which uses depth as an intermediate bridge to progressively align audio cues with visual features. Compared to state-of-the-art methods, DGCM-AVS achieves relative improvements of 10.2 percent in M_J and 8.7 percent in M_F on the AVSS dataset. We believe our study highlights depth as a promising yet underexplored modality for AVS and may encourage further research in this direction.
发表机构
- School of Intelligent Science and Technology, University of Science and Technology Beijing(北京科技大学智能科学与技术学院)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。