arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoViT:用于视觉Transformer的实例对应对比学习

CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer

Yisen Wang, Zhirong Wu, Limin Wang

arXiv 2609.01787首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology; Nanjing University(计算机软件新技术国家重点实验室; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉Transformer(ViT)无法区分目标实例的问题,提出CoViT框架,通过几何引导的对比学习为ViT注入实例感知,在多个实例级任务上实现超2AP点的稳定性能提升,且无需额外解码器或标签。

AI 中文摘要

视觉Transformer(ViT)在语义理解方面表现出色,但无法区分不同的目标实例(例如,两只狗会产生相同的嵌入),这限制了其在目标检测、实例分割等实例级任务中的应用。我们提出了对比视觉Transformer(CoViT),这是一种通过几何引导的对比学习为ViT注入实例感知能力的自监督学习框架。CoViT通过构建三元组独特地协调ViT的注意力图和嵌入:(1)注意力引导的掩码:通过自适应阈值和形态学操作优化多头注意力,生成实例掩码以识别前景锚点;(2)最难对比对挖掘:对于每个锚点,计算成对嵌入相似度以选择实例内最难正样本(其掩码内最不相似的补丁)和实例间最难负样本(来自其他实例的最相似补丁,负样本搜索期间会屏蔽实例内区域)。这些三元组驱动的对比损失同时压缩实例内方差并扩大实例间间隔,迫使ViT区分实例间细微的几何和外观差异。以ViT为骨干架构时,CoViT在多个实例级感知任务上始终实现超过2个AP点的稳定性能提升。值得注意的是,CoViT无需额外的解码器或标签,证明纯ViT可通过固有的注意力先验和针对性的对比约束学习实例感知表示。代码和模型将被发布。

英文摘要

Vision Transformers (ViT) excel in semantic understanding but fail to discriminate between object instances (e.g., identical embeddings for two dogs), limiting their use in instance-level tasks such as object detection and instance segmentation. We propose Contrastive Vision Transformer (CoViT), a self-supervised learning framework that injects instance-awareness into ViT through geometry-guided contrastive learning. CoViT uniquely coordinates ViT's attention maps and embeddings by constructing triplets: (1) Attention-guided masking: Refine multi-head attention via adaptive thresholding and morphological operations to generate instance masks, identifying foreground anchors; (2) Hardest contrastive mining: For each anchor, computing pairwise embedding similarities to select the intra-instance hardest positive (least similar patch within its mask) and inter-instance hardest negative (most similar patch from other instances), with intra-instance regions masked during negative search. These triplets drive a contrastive loss that simultaneously compresses intra-instance variance and expands inter-instance margins, forcing ViT to discern subtle geometric and appearance differences between instances. CoViT consistently achieves stable performance gains of over 2 AP points across multiple instance-level perception tasks by using ViT as backbone architecture. Notably, CoViT requires no extra decoders or labels, demonstrating that a pure ViT can learn instance-aware representations via inherent attention priors and targeted contrastive constraints. Code and models will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑