arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26600cs.CV

JEPADepth:用于自监督单目深度估计的掩码预测表示学习

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

Ionuţ Grigore, Călin-Adrian Popa

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出JEPADepth框架,结合I-JEPA的掩码预测损失优化自监督单目深度估计,在KITTI、Make3D等数据集上实现性能提升,且推理无额外成本。

中文摘要 AI 辅助

自监督单目深度估计通常依赖光度重构损失,该损失耦合了深度、位姿和外观假设。本文提出JEPADepth,这是一个受图像联合嵌入预测架构(I-JEPA)启发、融合互补训练目标的自监督单目深度框架,用于自监督深度学习。我们的方法在标准光度流程基础上,加入了在预训练DINOv3视觉Transformer编码器表示空间中计算的掩码预测损失:预测器在结构化掩码下从可见上下文区域嵌入推断目标区域嵌入,推理时与目标编码器一同丢弃,不增加部署成本。在KITTI数据集上,加入JEPA目标后,性能较基于DINOv3的相同光度基线持续提升,且未改变推理时架构。与现有单目自监督方法相比,JEPADepth与基于Transformer的最先进方法竞争力相当,在标准基准上优于强大的基于CNN的基线;在零样本迁移(在KITTI上训练且不微调直接评估)中,JEPADepth在Make3D和Cityscapes两个数据集的多项指标上,均达到对比方法中的最佳或接近最佳性能。

英文摘要

Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Predictive Architectures (I-JEPA) for self-supervised depth learning. Our method augments a standard photometric pipeline with a masked prediction loss computed in the representation space of a pretrained DINOv3 Vision Transformer encoder. A predictor infers target-region embeddings from visible context-region embeddings under structured masking, and is discarded along with the target encoder at inference time, adding no deployment cost. On KITTI, adding the JEPA objective consistently improves performance over the same DINOv3-based photometric baseline, without changing the inference-time architecture. Compared to prior monocular self-supervised methods, JEPADepth is competitive with state-of-the-art transformer-based approaches and outperforms strong CNN-based baselines on the standard benchmark. In zero-shot transfer (trained on KITTI and evaluated without fine-tuning), JEPADepth achieves the best or near-best performance among the compared methods on both Make3D and Cityscapes across multiple metrics.

发表机构

  • Politehnica University of Timișoara(蒂米什瓦拉理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑