深度假设引导的迭代细化用于事件-图像单目深度估计
Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HypoDepth,首个事件-图像单目深度迭代细化框架,通过离散深度假设体和三维代价体将深度回归转为约束搜索,在DSEC和MVSEC上取得最先进结果,且微型模型支持实时运行。
AI中文摘要:
事件相机具有优异的动态特性,在单目深度估计(MDE)方面展现出巨大潜力。然而,现有方法主要通过优化上下文特征来提升性能,但仍难以应对直接全深度回归的病态和非线性问题。本文提出HypoDepth,这是首个事件-图像单目深度迭代细化框架。通过引入离散的深度假设体(DHV),我们将深度回归问题转化为约束深度搜索任务。具体而言,我们在DHV特征与上下文特征之间构建三维代价体,并进行多尺度相关性搜索以引导稳定的残差优化。这种轻量级代价体能够在多分辨率下实现高效的从全局到局部的细化。我们的方法在DSEC和MVSEC数据集上以最先进的性能和强大的零样本泛化能力超越了现有方法。同时,我们的微型模型在精度与效率之间实现了出色平衡,能够在资源受限设备上实现实时性能。
英文摘要:
Event cameras hold excellent dynamic properties, showing great potential for monocular depth estimation (MDE). However, existing methods mainly improve performance by optimizing contextual features, but still struggle with the ill-posed and nonlinear nature of direct full-depth regression. In this paper, we propose HypoDepth, the first event-image monocular depth iterative refinement framework. By introducing a discrete Depth Hypothesis Volume (DHV), we transform the depth regression problem into a constrained depth search task. Specifically, we construct a 3D cost volume between the DHV features and contextual features and perform a multi-scale correlation search to guide stable residual optimization. This lightweight cost volume enables efficient global-to-local refinement across multi-resolution. Our method outperforms existing approaches on DSEC and MVSEC with state-of-the-art results and strong zero-shot generalization. Meanwhile, our tiny model achieves an excellent balance between accuracy and efficiency, enabling real-time performance on resource-limited devices.