arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30233cs.CV

面向通用视觉 grounding 的语义-空间判别性增强

Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding

Kaiyan Lei, Xu-Yao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对通用视觉 grounding 任务中易混淆相似目标的问题,提出 SSDE 框架,含 SeDE 与 SpDE 模块,在十个数据集上取得优异性能。

中文摘要 AI 辅助

通用视觉 grounding(Generalized Visual Grounding, GVG)任务旨在根据指代表达式在图像中定位目标,它通过整合多目标与非目标场景扩展了经典的视觉 grounding 范式。现有方法通常依赖全局语义匹配或粗粒度区域交互进行定位,其判别线索主要来自句子级语义或区域上下文。在复杂的多目标场景中,这类方法容易混淆视觉相似的目标,难以建立稳定的实例级决策边界。为解决这些局限,本文提出了一种用于通用视觉 grounding 的新型语义-空间判别性增强(Semantic-Spatial Discriminability Enhancement, SSDE)框架,旨在增强细粒度语义与空间定位的判别能力,提升跨模态理解与实例级 grounding 性能。具体而言,为增强查询表示在细粒度层面的语义判别性,我们提出了语义判别性增强(Semantic Discriminability Enhancement, SeDE)模块,该模块利用空间引导的交叉注意力解耦与目标相关的细粒度视觉属性,并将其与文本主语语义整合;此外,为增强被指代表目标的空间判别性,我们引入了空间判别性增强(Spatial Discriminability Enhancement, SpDE)模块,该模块建模实例中心密度图以表征目标的空间分布,并通过将其作为辅助监督信号在空间域显式构建实例分离结构。大量实验表明,SSDE 在经典及通用视觉 grounding 任务的共十个数据集上均取得了优异性能。

英文摘要

Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous methods typically rely on global semantic matching or coarse-grained region interactions for localization, where the discriminative cues are primarily derived from sentence-level semantics or regional context. In complex multi-target scenarios, such approaches tend to confuse visually similar targets, making it difficult to establish stable instance-level decision boundaries. To address these limitations, this paper proposes a novel Semantic-Spatial Discriminability Enhancement (SSDE) framework for generalized visual grounding, which aims to enhance the discriminative ability on fine-grained semantics and spatial localization, improving both cross-modal understanding and instance-level grounding. Specifically, to enhance the semantic discriminability of query representations at the fine-grained level, we propose a Semantic Discriminability Enhancement (SeDE) module, which leverages spatially guided cross-attention to disentangle fine-grained target-relevant visual attributes and integrates them with the textual subject semantics. Furthermore, to strengthen the spatial discriminability of the referred targets, we introduce a Spatial Discriminability Enhancement (SpDE) module, which models an instance center density map to characterize the spatial distribution of targets, and explicitly constructs instance separation structures in the spatial domain by employing them as an auxiliary supervision signal. Extensive experiments show that SSDE achieves superior performance on ten datasets across both classic and generalized visual grounding tasks.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑