arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceCLIP:从Patch-to-CLS贡献中恢复局部语义

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Xinran Liu, Shouqian Shi, Yutong Chen, Ge Wang, Xin-Wei Yao, Sheng Zhong

arXiv 2607.26107首次发表:更新:

AI 中文总结

TraceCLIP是一种无训练视觉-语言框架,通过分离CLS注意力输出的Patch特定术语恢复局部语义,在8个零样本语义分割基准上较现有无训练方法平均mIoU提升1.3至4.5个点。

AI 中文摘要

密集视觉-语言理解(包括目标定位、区域识别和开放词汇语义分割)需要将语言概念与空间定位的视觉区域关联起来。CLIP通过大规模对比预训练学习共享的图像-文本嵌入空间,为这些任务提供了坚实基础,但它的图像级目标是将文本与CLS生成的全局表示对齐,仅间接约束局部视觉-语言对应关系。现有方法要么引入额外监督、外部模型或特定任务适配,而无训练方法主要从现有Patch特征中恢复密集响应,未探究CLIP内部局部语义最易获取的位置。我们提出TraceCLIP,一种无训练框架,通过分离写入CLS注意力输出的Patch特定术语,恢复潜在的Patch级语义证据;TraceCLIP进一步将贡献衍生的语义响应转换为语义-测地线拓扑门,以校准最终层Patch亲和力用于密集特征重建。诊断实验表明,这些贡献特征表现出强局部语义判别力和文本条件空间对齐能力。在8个零样本语义分割基准上,TraceCLIP在无额外训练、外部视觉基础模型或区域级监督的情况下,相对于最强的现有无训练方法,在两种骨干网络和背景设置下,平均mIoU提升了1.3至4.5个点。更广泛地说,这些发现表明,空间定位的语义可能在全局对齐表示的内部结构中仍可获取。

英文摘要

Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.

Comments9 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑