arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynCLIP:用于鲁棒开放词汇密集感知的同义词一致语言-图像预训练

SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng

arXiv 2607.11008首次发表:更新:

发表机构

Beihang University; State Key Laboratory of Space Information System and Integrated Application; Nanjing University; Harbin Engineering University; Beijing Zhongguancun Academy(北京航空航天大学; 空间信息系统与集成应用国家重点实验室; 南京大学; 哈尔滨工程大学; 北京中关村科学城)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对开放词汇密集感知中同义词引起的定位不一致问题,提出SynCLIP框架,通过SSA和SAR模块增强注意力一致性与定位精度,构建SEViC支持预训练,实验证明该方法显著提升定位一致性并达领先性能。

AI 中文摘要

开放词汇密集感知(OVDP)旨在通过利用文本知识来定位训练期间未见过的物体。尽管基于CLIP的方法最近取得了显著进展,但存在一个关键限制:同义词引起的定位不一致,即语义等效的表达式会产生不同的空间注意力模式。这种不一致削弱了现有方法在实际OVDP应用中的鲁棒性和性能。为解决此问题,我们提出了SynCLIP,一个同义词一致语言-图像预训练框架,用于增强OVDP的同义词鲁棒定位。SynCLIP引入了语义一致空间注意力对齐(SSA)模块,通过最小化原始和同义词表达式的注意力图之间的差异来增强空间注意力一致性。此外,空间注意力细化(SAR)模块在对齐图中选择性地加强最语义相关的空间区域,以实现更精确和稳定的定位。为支持同义词一致预训练,我们还构建了一个同义词丰富视觉语料库(SEViC),用多个同义词和文本定义扩充每个类别。在多个基准上的广泛实验表明,SynCLIP在不同语言变体下显著提高了定位一致性,并在基于CLIP的OVDP方法中取得了领先性能。

英文摘要

Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynCLIP, a Synonym-Coherent Language-Image Pretraining framework that enhances synonym-robust grounding for OVDP. SynCLIP introduces a Semantic-consistent Spatial Attention alignment (SSA) module to enhance spatial attention consistency by minimizing discrepancies between attention maps of original and synonymous expressions. Furthermore, a Spatial Attention Refinement (SAR) module selectively strengthens the most semantically relevant spatial regions within aligned maps for more precise and stable grounding. To support synonym-coherent pretraining, we also construct a Synonym-Enriched Visual Corpus (SEViC), which augments each category with multiple synonyms and textual definitions. Extensive experiments on multiple benchmarks demonstrate that SynCLIP substantially improves grounding consistency under diverse linguistic variants and achieves state-of-the-art performance among CLIP-based OVDP methods. Code is available at https://github.com/Justlovesmile/SynCLIP.

CommentsAccepted by CVPR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑