arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17956cs.RO

鲁棒视觉惯性里程计需要更多学习吗?分布偏移下的几何验证视觉测量

Does Robust VIO Need More Learning? Geometry-Verified Visual Measurements under Distribution Shift

  • Tongji University(同济大学)
  • Nanjing University of Information Science and Technology(南京信息工程大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Sun Yat-sen University(中山大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yangyang Ning, Shu Liang, Quanbo Ge, Tianchen Deng, Yuhua Qi, Shenghai Yuan

AI总结:

研究分布偏移下鲁棒视觉惯性里程计是否需更多学习,提出最小学习立体VIO框架,将学习限于视觉测量生成,结合几何验证等,实验表明该方法在多种场景下准确稳定,精心集成学习视觉测量更有效。

AI中文摘要:

学习越来越多地被引入视觉惯性里程计(VIO),从学习特征前端到以学习为主的运动和几何估计。然而,当部署条件与训练分布不同时,更多的学习不一定能提高鲁棒性。本文探讨分布偏移下的鲁棒VIO是否真的需要更深层次的学习估计,还是可以将学习局限于视觉测量生成。提出了一个最小学习立体VIO框架,其中SEA-RAFT仅用于提出密集立体对应并预测其不确定性,而时间跟踪、几何验证和状态估计保持显式。在稀疏特征位置采样密集流,使用预测的不确定性和立体极线一致性进行滤波,并通过不确定性加权重投影因子纳入滑动窗口立体惯性估计器。相同的不确定性通过立体三角测量进一步传播用于下游各向异性3D高斯映射。在EuRoC、VIODE和4Seasons上的实验表明,在运动模糊、动态场景、光照变化和从室内到室外的大分布偏移下,能进行准确稳定的估计。消融实验表明,仅靠学习流是不够的:收益来自于将学习到的对应提议与几何验证和不确定性感知加权相结合。这些结果表明,对于OOD鲁棒VIO,精心集成的学习视觉测量可能比学习更大比例的估计管道更有效。基准测试的代码和配置将在接受后开源。可在这个https URL获取补充视频。

英文摘要:

Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-ends to learning-dominant motion and geometry estimation. However, learning more of the pipeline does not necessarily improve robustness when deployment conditions differ from the training distribution. This work asks whether robust VIO under distribution shift truly requires deeper learned estimation, or whether learning can be confined to visual measurement generation. We propose a minimal-learning stereo VIO framework in which SEA-RAFT is used only to propose dense stereo correspondences and predict their uncertainty, while temporal tracking, geometric verification, and state estimation remain explicit. Dense flow is sampled at sparse feature locations, filtered using predicted uncertainty and stereo epipolar consistency, and incorporated into a sliding-window stereo-inertial estimator through uncertainty-weighted reprojection factors. The same uncertainty is further propagated through stereo triangulation for downstream anisotropic 3D Gaussian mapping. Experiments on EuRoC, VIODE, and 4Seasons demonstrate accurate and stable estimation under motion blur, dynamic scenes, illumination changes, and large indoor-to-outdoor distribution shifts. Ablations show that learned flow alone is insufficient: the gains arise from combining learned correspondence proposals with geometric verification and uncertainty-aware weighting. These results suggest that, for OOD-robust VIO, carefully integrated learned visual measurements can be more effective than learning a larger fraction of the estimation pipeline. Code and configs for the benchmark will be open-source upon acceptance. A supplementary video is available at https://drive.google.com/file/d/1EVRhOkhanmNXHbQS1Vr80FoEIAYOYOV2/view

↑