学习先验何时有助于视觉惯性估计?关于先验集成、标定、初始化和后端一致性的受控研究
When Do Learned Priors Help Visual Inertial Estimation? A Controlled Study of Prior Integration, Calibration, Initialization, and Backend Consistency
- Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过受控框架研究学习先验在视觉惯性估计中的价值,发现融合性能提升可能源于后端或标定变化,而非先验本身,强调需同后端对照和联合分析才能可靠评估。
AI中文摘要:
学习组件日益被集成到几何视觉惯性估计器中,以提供运动、深度、偏置、不确定性或置信度线索。然而,尚不清楚性能提升是源于有用的学习先验,还是源于后端、标定、初始化、时间关联或评估基准的变化。我们提出了一个用于学习增强视觉惯性估计的受控框架,该框架将融合增益与学习先验的增量价值分离开来,并评估四个证据层:局部运动一致性、全局轨迹精度、物理状态正确性和数值一致性。我们使用基于MonoViT的单目运动先验实例化该框架,将其作为局部相对运动因子添加到未改变的VINS后端。我们在匹配的传感器流、时间戳、初始化、前端/后端设置和相机-IMU外参下比较原始VINS和学习先验VINS,同时探测标定、初始化、状态耦合、尺度、束调整和回环闭合。在KITTI上,使用固定参考外参,原始VINS的平移APE RMSE为31.4米,而使用学习先验的为31.8米。在四次记录中,先验仅使平均APE变化-0.2%,而五倍高的权重使其恶化8.2%。在线外参更新分别使平均APE增加45.7%和52.1%,而平均RPE变化小于2%。这些结果表明,仅融合性能不能确立学习先验的价值。可靠的评估需要同后端对照,以及对先验兼容性、标定、初始化、全局漂移、物理状态误差和后端一致性的联合分析。
英文摘要:
Learned components are increasingly integrated into geometric visual--inertial estimators to provide motion, depth, bias, uncertainty, or confidence cues. Yet it remains unclear whether gains arise from useful learned priors or from changes in the backend, calibration, initialization, temporal association, or evaluation gauge. We present a controlled framework for learning-augmented visual--inertial estimation that separates fusion gain from the incremental value of a learned prior and evaluates four evidence layers: local motion consistency, global trajectory accuracy, physical-state correctness, and numerical consistency. We instantiate the framework with a MonoViT-based monocular motion prior added as a local relative-motion factor to an unchanged VINS backend. We compare Original VINS and learned-prior VINS under matched sensor streams, timestamps, initialization, frontend/backend settings, and camera--IMU extrinsics, while probing calibration, initialization, state coupling, scale, bundle adjustment, and loop closure. On KITTI, with fixed reference extrinsics, translation APE RMSE is 31.4 m for Original VINS and 31.8 m with the learned prior. Across four recordings, the prior changes mean APE by only -0.2%, while a five-times-higher weight worsens it by 8.2%. Online extrinsic updates increase mean APE by 45.7% and 52.1%, respectively, while mean RPE changes by less than 2%. These results show that fusion performance alone cannot establish the value of learned priors. Reliable evaluation requires same-backend controls and joint analysis of prior compatibility, calibration, initialization, global drift, physical-state error, and backend consistency.