发表机构
Hefei University of Technology; Cardiff University; University of Macau(合肥工业大学; 卡迪夫大学; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放词汇视听语义分割的现有方法缺陷,提出AGCL框架及AMCG、AGTA、SDM模块,在AVSBench-OV数据集上显著优于现有SOTA方法,尤其在未见类别上表现突出。
AI 中文摘要
开放词汇视听语义分割(OV-AVSS)旨在对开放类别集合中发声物体进行像素级分割。现有方法依赖类别无关的前景定义,将语义多样的物体归为异质正集,导致模型学习到不稳定的发声模式并生成不可靠的候选区域。为解决该问题,我们将目标重新表述为类别特定的,并提出新型声学接地代价学习(AGCL)框架,将静态的、与音频无关的视觉-文本先验转换为动态的、音频接地的代价表示。对于类别内发声性发现,我们设计音频调制代价生成(AMCG)模块与音频引导时间聚合(AGTA)模块,通过低侵入性的音频注入机制实现帧级发声区域突出和视频级时间优化。对于类别间干扰项区分,我们引入协同干扰项挖掘(SDM)策略,选择性惩罚声学和语义混淆的负类别以学习更具判别性的决策边界。在AVSBench-OV数据集上的大量实验表明,我们的方法显著优于现有最先进方法,尤其在未见类别上表现突出。代码可在指定URL获取。
英文摘要
Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method relies on a class-agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category-specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio-agnostic visual-text priors into dynamic, audio-grounded cost representations. For intra-category soundingness discovery, we devise Audio-Modulated Cost Generation (AMCG) and Audio-Guided Temporal Aggregation (AGTA) modules to enable both frame-level sounding region highlighting and video-level temporal refinement with a low-intrusive audio injection mechanism. For inter-category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench-OV dataset demonstrate that our method significantly outperforms previous state-of-the-art approaches, particularly on unseen categories. Code is available at https://github.com/spyflying/AGCL.
CommentsAccepted by ACM MM 2026