DAPEVO:深度自适应补丁帧-事件视觉里程计
DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry
查看机构详情
- University of Zurich(苏黎世大学)
- Bocconi University(博科尼大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
DAPEVO是一种深度自适应补丁帧-事件视觉里程计,通过独立估计图像与事件对应并融合相关性证据,在RGB帧稀疏或退化时显著降低轨迹误差,优于现有方法。
中文摘要 AI 辅助
视觉里程计对于GPS受限环境中的自主导航至关重要,然而基于RGB的方法仍然容易受到运动模糊、具有挑战性的光照和帧丢失的影响。事件相机以高时间分辨率和动态范围补充了传统相机,但其异步测量使可靠对应估计变得复杂。我们提出了DAPEVO,一种学习型视觉里程计系统,它在共享补丁位置独立估计图像和事件对应关系,并在运动细化之前融合它们的相关性证据。每个跟踪的补丁维护图像和事件描述符,一个学习型标量门在共享循环细化和束调整更新之前,为每个补丁-帧边组合模态特定的相关性嵌入。DAPEVO还支持仅事件观测,当RGB帧稀疏或不可用时能够继续跟踪,而模态感知的关键帧剔除保留了稀缺的帧约束。在UZH-FPV上,当仅保留六分之一的RGB帧时,DAPEVO的平均绝对轨迹误差(ATE)仅增加36%,从1.00米增加到1.36米,而DPVO和RAMP-VO的ATE分别增加了3.7倍和3.1倍。在TartanEvent上,DAPEVO在3Hz RGB输入下同样保持低于1米的ATE,而DPVO和RAMP-VO超过9米。在TartanEvent上RGB输入退化的情况下,DAPEVO实现了0.60米的ATE,而DPVO和RAMP-VO均超过4米,同时也优于仅事件的DEVO(0.87米)。
英文摘要
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch locations and fuses their correlation evidence before motion refinement. Each tracked patch maintains image and event descriptors, and a learned scalar gate combines modality-specific correlation embeddings for each patch--frame edge before a shared recurrent refinement and bundle-adjustment update. DAPEVO also supports event-only observations, enabling continued tracking when RGB frames are sparse or unavailable, while modality-aware keyframe culling preserves scarce frame constraints. On UZH-FPV, when retaining only one in six RGB frames, DAPEVO's mean absolute trajectory error (ATE) increases by only 36%, from 1.00 to 1.36m, whereas the ATE of DPVO and RAMP-VO rises by factors of $3.7\times$ and $3.1\times$, respectively. On TartanEvent, DAPEVO similarly remains below 1m ATE at 3Hz RGB input, while DPVO and RAMP-VO exceed 9m. Under degraded RGB input on TartanEvent, DAPEVO achieves an ATE of 0.60m, compared with more than 4m for both DPVO and RAMP-VO, while also outperforming event-only DEVO at 0.87m.