arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24125cs.CVcs.CG

正样本对几何至关重要:用于视觉表征对比学习的最优传输

Positive Pair Geometry Matters: Optimal Transport for Contrastive Learning of Visual Representations

Akshit Nanda, Shahzad Ahmad, Ram Prasad Padhy

首次发表
浏览论文内容

中文总结 AI 辅助

针对随机增强可能破坏语义的问题,提出基于熵最优传输位移插值构造几何一致正样本的OTCLR框架,在不改编码器下提升表征质量与迁移性能。

中文摘要 AI 辅助

对比自监督学习通过从同一图像的多重增强视图学习表征,已取得强劲性能。然而,现有大多数方法使用独立采样的随机增强来构造正样本对,这可能改变语义内容并忽略数据分布的内在几何结构。在本工作中,我们提出OTCLR,一种最优传输感知的对比学习表征框架,用于生成几何一致的正样本。我们不直接对比两个随机增强视图,而是通过熵最优传输位移插值,在原始图像与其增强变体之间构造中间视图。这些传输插值样本作为正视图,能更好地保留图像结构,同时显式建模空间分布几何。为进一步促进平滑的表征学习,我们评估辅助的Sinkhorn正则化项,以鼓励传输插值视图与其端点图像保持一致。所提方法可无缝融入标准对比学习流程,无需修改编码器架构。在多个基准数据集上的实验表明,与传统基于增强的对比学习基线相比,我们的方法提升了表征质量和迁移学习性能。

英文摘要

Contrastive self-supervised learning has achieved strong performance by learning representations from multiple augmented views of the same image. However, most existing methods construct positive pairs using independently sampled stochastic augmentations, which may alter semantic content and ignore the intrinsic geometry of the data distribution. In this work, we propose OTCLR, an optimal transport-aware framework for contrastive learning representations that generates geometry-consistent positive samples. Instead of directly contrasting two randomly augmented views, we construct intermediate views between the original image and its augmented variants through entropic optimal-transport displacement interpolation. These transport-interpolated samples serve as positive views that better preserve image structure while explicitly modeling spatial distributional geometry. To further promote smooth representation learning, we evaluate auxiliary Sinkhorn regularization terms that encourage transport-interpolated views to remain consistent with their endpoint images. The proposed method can be incorporated into standard contrastive learning pipelines without modifying the encoder architecture. Experiments on multiple benchmark datasets show that our approach improves representation quality and transfer learning performance compared with conventional augmentation-based contrastive learning baselines.

发表机构

  • Indian Institute of Technology Bhubaneswar(印度理工学院布巴内斯瓦尔分校)
  • Østfold University of Applied Sciences(东福尔应用科学大学)
  • NTNU(挪威科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑