FFVO:用于长程视觉里程计的前馈位姿解码器
FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry
查看机构详情
- Waymo LLC(Waymo 有限责任公司)
- University of Washington(华盛顿大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
FFVO提出一种前馈位姿解码器,通过紧凑令牌表示、分层时序解码和轨迹监督,在长程视觉里程计中实现高效、稳定且低漂移的相机位姿估计。
中文摘要 AI 辅助
稳定可靠的4D空间理解是自动驾驶系统的基础。虽然前馈重建网络可以一次性估计相机运动和3D结构,但在长视频上的位姿估计仍面临计算成本、长上下文模糊性和时间不稳定性的挑战。为解决这些问题,我们提出了前馈视觉里程计(FFVO),一种针对联合重建架构的位姿特化适配,用于高效且时间稳定的相机位姿估计。FFVO采用(i)紧凑的相机令牌表示,用于计算高效的时序聚合;(ii)分层的局部到全局时序解码器,通过分离短程运动聚合与序列级集成来缓解几何模糊性;(iii)中间轨迹监督,以促进时间稳定性。在Waymo开放数据集(WOD)、KITTI以及大规模专有基准上的广泛评估表明,我们的方法优于现有前馈方法,并大幅减少了抖动和漂移。这些结果支持FFVO作为长程视觉里程计设置中有效的相机位姿前馈解码器。
英文摘要
Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate camera motion and 3D structure in one pass, pose estimation over long videos remains challenged by computational cost, long-context ambiguity, and temporal instability. To address these challenges, we propose Feedforward Visual Odometry (FFVO), a pose-specialized adaptation of joint reconstruction architectures for efficient and temporally stable camera-pose estimation. FFVO uses (i) a compact camera-token representation for computationally efficient temporal aggregation, (ii) a hierarchical local-to-global temporal decoder that mitigates geometric ambiguity by separating short-range motion aggregation from sequence-level integration, and (iii) intermediate trajectory supervision that promotes temporal stability. Extensive evaluation on the Waymo Open Dataset (WOD), KITTI, and a large-scale proprietary benchmark demonstrates that our method performs favorably against existing feedforward approaches, and greatly reduces jitter and drift. These results support FFVO as an effective feedforward camera-pose decoder in long-horizon visual odometry settings.