深度引导的对比学习:用于具有3D空间感知的2D表示
Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness
- KU Leuven(鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出深度引导的对比学习(DGCL),利用相对3D距离比较将空间感知注入2D表示,提升场景理解与下游语义任务性能。
AI中文摘要:
标准的对比学习框架主要从语义角度设计,然而学习保留3D空间结构的2D视觉表示对于场景理解同样重要。在这项工作中,我们提出了深度引导的对比学习(DGCL),这是一种简单的辅助目标,将3D空间感知注入到2D对比表示学习中。我们的关键思想是利用深度将局部3D邻近性转化为对比相似性:在3D空间中更近的像素被鼓励拥有比相距更远的像素更相似的表示。DGCL不依赖于绝对深度值,而是通过随机采样像素之间的相对3D距离比较来制定监督,使得目标对深度尺度不变、计算高效,并且易于集成到现有的对比学习框架中。跨不同数据集和模型的实验表明,DGCL持续改进2D表示学习,并通过更强的空间和几何理解有益于语义下游任务。代码可在 https://github.com/LeungTsang/DGCL 获取。
英文摘要:
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objective that injects 3D spatial awareness into 2D contrastive representation learning. Our key idea is to use depth to convert local 3D proximity into contrastive similarity: pixels that are closer in 3D space are encouraged to have more similar representations than pixels that are farther apart. Instead of relying on absolute depth values, DGCL formulates supervision through relative 3D distance comparisons among randomly sampled pixels, making the objective invariant to depth scale, efficient to compute, and easy to integrate into existing contrastive frameworks. Experiments across different datasets and models show that DGCL consistently improves 2D representation learning and benefits semantic downstream tasks by stronger spatial and geometric understanding. The code is available on https://github.com/LeungTsang/DGCL.