AspectCLIP:通过方面引导的一致性正则化优化CLIP表示空间
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
浏览论文内容
中文总结 AI 辅助
研究针对图像与文本信息不对称问题,提出AspectCLIP框架,通过基于文本相似性划分样本为属性簇,在簇内应用循环一致性并限制跨簇正则化,实现更优的表示空间,在下游任务中性能超传统方法。
中文摘要 AI 辅助
对比语言-图像预训练通过大规模对比学习学习共享表示空间。然而,现有的强制全局一致性正则化的方法忽略了一个关键挑战:图像和文本之间固有的信息不对称,字幕通常只描述图像的一个特定方面,因此具有相似视觉内容的图像可以与完全不同的文本内容和语义信息配对。结果,全局正则化器在字幕描述不同方面的视觉相似图像之间无意中施加了约束,在表示空间中引入了语义失真。我们提出了AspectCLIP,一个重新制定一致性正则化以尊重这种一对多结构的框架。AspectCLIP首先根据文本相似性将训练样本划分为属性簇以识别方面一致的组,然后在每个簇内应用完全循环一致性,同时将跨簇正则化限制在原型级比较。这种方面引导的正则化仅在图像和文本描述一致的方面时才强制严格的几何对齐,同时允许在不同方面之间具有灵活性。在下游任务上的广泛实验表明,AspectCLIP始终优于传统方法,并实现了更结构化的表示空间。
英文摘要
Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the inherent information asymmetry between images and text: captions typically describe only one specific aspect of an image, thus images with similar visual content can be paired with completely divergent textual content and semantic information. Consequently, global regularizers inadvertently impose constraints between visually similar images whose captions describe divergent aspects, introducing semantic distortion into the representation space. We propose AspectCLIP, a framework that reformulates consistency regularization to respect this one-to-many structure. AspectCLIP first partitions training samples into attribute clusters based on textual similarity to identify aspect-coherent groups, then applies full cyclic consistency within each cluster while restricting cross-cluster regularization to prototype-level comparisons. This aspect-guided regularization enforces strict geometric alignment only when images and texts describe a consistent facet, while allowing flexibility across divergent aspects. Extensive experiments on downstream tasks demonstrate that AspectCLIP consistently outperforms traditional methods and achieves a more structured representation space.
发表机构
- School of Computer Science and Technology, South China University of Technology(华南理工大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。