EpiTransfer:基于时间单目航空帧的稀疏、免训练远距离深度估计
EpiTransfer: Sparse, Training-Free Long-Range Depth Estimation from Temporal Monocular Aerial Frames
浏览论文内容
中文总结 AI 辅助
本文提出一种免训练的极线转移深度估计方法,利用两幅单目图像和位姿合成虚拟立体对,在室内外场景中达到与直接三角化相当的精度,并显著优于未微调的学习式基线。
中文摘要 AI 辅助
可靠的3D空间理解对于自主导航、避障和场景重建至关重要。虽然最先进的学习式深度估计技术在分布内数据上能达到高精度,但它们通常难以泛化到新颖的视角和高度。本文提出了一种基于几何推导的免训练深度估计方法,利用极线转移,仅需两张单目图像和相机位姿估计。通过利用相机运动合成具有自由选择基线的虚拟立体对,我们的方法将时间对应关系转化为立体三角化任务,同时缓解了直接双视角三角化固有的几何退化问题。在户外无人机飞行(最大范围约90米)和室内OptiTrack环境中,以LiDAR地面真值进行验证,该方法在室内实现了0.092的绝对相对误差(AbsRel)和0.940的δ<1.25精度,与直接三角化(AbsRel 0.073)相当,同时在更具挑战性的场景中保留了更大比例的有效深度,并且大幅优于现成的基于学习的基线方法,如ZoeDepth(AbsRel 0.225)和Depth Anything V2(AbsRel 0.570),这些方法未针对该领域进行训练或微调,且我们的方法无需任何训练数据。
英文摘要
Reliable 3D spatial understanding is essential for autonomous navigation, obstacle avoidance, and scene reconstruction. While state-of-the-art learned depth estimation techniques achieve high accuracy in-distribution, they often generalize poorly to novel viewpoints and altitudes. This paper presents a geometrically derived, training-free depth estimation method using epipolar transfer with only two monocular images and camera pose estimates. By leveraging camera motion to synthesize a virtual stereo pair with a freely chosen baseline, our approach transforms temporal correspondence into a stereo triangulation task while mitigating geometric degeneracies inherent to direct two-view triangulation. Validated across outdoor drone flights (to a maximum range of approximately 90\,m) and indoor OptiTrack environments against LiDAR ground truth, the method achieves an indoor AbsRel of 0.092 and $δ< 1.25$ of 0.940, comparable to direct triangulation (AbsRel 0.073) while retaining valid depth over a larger fraction of challenging scenes, and substantially outperforms off-the-shelf learning-based baselines such as ZoeDepth (AbsRel 0.225) and Depth Anything V2 (AbsRel 0.570), which are not trained or fine-tuned for this domain, with no training data required.
发表机构
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。