发表机构
Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于核典型相关分析(KCCA)的3vKCCA方法,通过最大化特征子空间投影相关性来增强视觉语言模型的细粒度视觉感知,在MMVP-VLM上准确率从17.8提升至25.9,且保持零样本检索性能。
AI 中文摘要
视觉语言模型如CLIP展现出强大的语义泛化能力,但在细粒度视觉感知方面仍存在局限。近期工作KUEA提出了一种自然解决方案,通过在视觉中心DINOv2的监督下微调图像编码器,逐元素对齐其核矩阵,同时正则化嵌入以保持与预训练视觉编码器的接近,从而保留CLIP中的图像-文本语义。然而,我们表明,减弱与DINOv2的对齐损失的作用并不一定会降低其细粒度视觉性能,这表明核矩阵差异可能不足以用于进一步的视觉表示增强,这促使我们重新审视对齐公式。在这项工作中,我们提出了一种新的视角,通过核典型相关分析(KCCA)在特征子空间上表征表示对齐,该分析最大化投影相关性。在优化方面,我们基于KKT条件推导了一种高效的端到端训练方案,避免了KCCA中的特征值问题。进一步,我们将我们的方法扩展到三视图公式,即3vKCCA,其中预训练文本编码器的投影也被纳入统一优化框架中进行联合对齐。使用CLIP ViT-L/14在ImageNet-1K上,我们的3vKCCA将MMVP-VLM准确率从17.8提高到25.9,显著优于现有方法,同时保持了零样本图像-文本检索性能。
英文摘要
Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.