arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从校正立体中进行几何蒸馏:利用对极线索进行单目深度估计

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

Jung-Hee Kim, Xiaoming Liu

arXiv 2607.15600首次发表:更新:

发表机构

Michigan State University; University of North Carolina at Chapel Hill(密歇根州立大学; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对单目深度估计在不同环境下的难题,提出将多视图模型的尺度感知几何先验转移到单目深度基础模型的框架,通过对极蒸馏方法,无需多视图输入就能保持几何一致性,显著提升零样本度量深度估计性能,且与模型无关。

AI 中文摘要

单目深度基础模型在不同环境中展现出显著的泛化能力,但在不同环境下的度量深度估计仍存在困难,这源于单视图推理固有的尺度模糊性。近期多视图基础模型利用跨视图线索学习稳健的场景级几何和一致尺度,但在单图像推理时性能会下降。为弥合差距,我们提出一个框架,将多视图模型的尺度感知几何先验转移到单目深度基础模型中。具体介绍了对极蒸馏(EpiDistill),利用校正立体令牌,使单视图预测模型在推理时无需多视图输入就能保留对极注意力模式并保持几何一致性。实验结果表明,该方法显著改进了零样本度量深度估计,尤其在ETH3D和DIODE等具有挑战性的数据集上,且该方法与模型无关,能持续提升基于ViT的模型性能。

英文摘要

Monocular depth foundation models have demonstrated remarkable generalization capabilities across diverse environments. However, they continue to struggle with metric depth estimation in diverse environments. This limitation stems from the inherent scale ambiguity of single-view inference, leading to misaligned scale predictions even when the relative geometry is accurate. Conversely, recent multi-view foundation models leverage cross-view cues to learn robust scene-level geometry and consistent scale. Yet, these benefits typically vanish during single-image inference, as the absence of explicit geometric constraints causes performance to degrade. To bridge this gap, we propose a novel framework that transfers the scale-aware geometric priors of multi-view models into monocular depth foundation models. Specifically, we introduce an Epipolar Distillation (EpiDistill), an approach utilizing Rectified Stereo Tokens, which enables the single-view prediction model to retain epipolar attention patterns and maintain geometric consistency without requiring multi-view inputs at inference. Experimental results demonstrate that our method significantly improves zero-shot metric depth estimation, particularly on challenging datasets like ETH3D and DIODE where scale alignment is critical. Furthermore, our approach is model-agnostic, consistently boosting the performance of state-of-the-art ViT-based models, including UniDepthV2 and DepthPro.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑