arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAC-I$^2$:学习型度量感知协方差用于初始化与标定中的鲁棒视觉-惯性融合

MAC-I$^2$: Learned Metrics-Aware Covariance for Robust Visual-Inertial Fusion in Initialization and Calibration

Xiang Fei, Yuheng Qiu, Can Xu, Yutian Chen, Ruogu Li, Xingxing Zuo, Wenshan Wang, Sebastian Scherer

arXiv 2609.07116首次发表:更新:

发表机构

The Robotics Institute, Carnegie Mellon University; University of Toronto Institute for Aerospace Studies (UTIAS), University of Toronto; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(卡内基梅隆大学机器人研究所; 多伦多大学航空航天研究所(多伦多大学); 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MAC-I$^2$,通过学习度量感知协方差实现鲁棒视觉-惯性融合,在初始化与标定中显著提升成功率并降低误差。

AI 中文摘要

视觉-惯性(VI)融合是准确且鲁棒的状态估计的基础,其中相机和IMU测量根据各自的不确定性进行组合。然而,现有方法使用预定义的不确定性来融合这两种模态,而不考虑每种模态在局部上下文中的可靠性,因此在涉及光照变化、动态物体和无纹理区域的挑战性环境中常常表现不佳。在本文中,我们提出了MAC-I$^2$,通过为两种模态学习度量感知协方差来实现鲁棒的VI融合,使得视觉和IMU根据自身优势进行竞争,而不是依赖预定义的不确定性。这里,度量感知意味着每个预测的协方差忠实反映相应测量噪声的实际大小。在视觉方面,我们将学习到的特征匹配不确定性传播到位姿协方差以用于融合。在惯性方面,受积分误差在早期阶段急剧累积而后缓慢增长这一观察的启发,我们设计了一个具有可学习初始协方差的学习型IMU模型,并提出了一种在保留训练子集上的专门微调策略,以在未见序列上实现度量感知协方差。作为展示,我们构建了一个VI初始化和标定系统,因为准确且鲁棒的初始化和标定是任何可靠VI系统的前提。在EuRoC和VBR上的实验表明,MAC-I$^2$显著优于现有方法:它在EuRoC上实现了99.9%的初始化成功率,将重力和速度误差相比最强基线分别降低了约60%和42%,并在具有挑战性的VBR序列上保持了80%的成功率,而VINS-Mono等基线方法在这些序列上的成功率降至10%以下。

英文摘要

Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measurements are combined according to their respective uncertainties. Existing methods, however, fuse the two modalities with predefined uncertainties, regardless of how reliable each is in the local context, and thus often struggle under challenging environments involving illumination changes, dynamic objects, and textureless regions. In this paper, we present MAC-I$^2$, which achieves robust VI fusion through learned metric-aware covariance for both modalities, so that vision and IMU compete on their own merits rather than relying on predefined uncertainties. Here, metrics-aware means that each predicted covariance faithfully reflects the actual magnitude of the corresponding measurement noise. On the visual side, we propagate learned feature-matching uncertainties into pose covariances for the fusion. On the inertial side, motivated by the observation that integration error accumulates sharply at the early stage and grows slowly afterward, we design a learned IMU model with a learnable initial covariance, and propose a dedicated fine-tuning strategy on a held-out training subset to enable the metrics-aware covariance on unseen sequences. As a showcase, we build a VI initialization and calibration system, since accurate and robust initialization and calibration are the prerequisite for any reliable VI system. Experiments on EuRoC, and VBR show that MAC-I$^2$ substantially outperforms existing methods: it achieves a 99.9% initialization success rate on EuRoC, reducing gravity and velocity errors by about 60% and 42% over the strongest baseline, and maintains 80% success rate on challenging VBR sequences where baseline methods such as VINS-Mono drop below 10%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑