发表机构
Nanyang Technological University; Dalian University of Technology; Queensland University; Sun Yat-sen University(南洋理工大学; 大连理工大学; 昆士兰大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对地下环境VIO在传感器退化、误标定和动态遮挡下的失效问题,提出基于CERBERUS数据集的失效中心基准,系统评估四种VIO系统在九种扰动下的鲁棒性,揭示范式间脆弱性差异并提供部署指导。
AI 中文摘要
视觉惯性里程计(VIO)是GPS受限地下环境中自主运行的核心能力,然而在传感器漂移、标定误差和动态遮挡条件下,其可靠性可能急剧下降。现有评估主要强调标称条件下的精度,对实际部署失效何时发生提供的洞察有限。在本工作中,我们利用CERBERUS数据集,针对地下环境中的VIO提出了一种以失效为中心的应力测试基准。我们系统评估了四种具有代表性的VIO系统,涵盖滤波、优化和学习三种范式,在九种实际扰动设置下进行测试,包括IMU偏置与噪声变化、相机内参和外参漂移以及动态场景遮挡。除传统轨迹误差外,我们通过覆盖率(coverage ratio)和失效阈值(failure thresholds)分析鲁棒性极限,揭示了仅凭标称条件性能无法捕捉的失效行为。我们的研究显示了不同VIO范式间的不同脆弱性模式:一些方法对惯性退化更敏感,而另一些则更受几何误标定或动态干扰的影响。这些结果为VIO选择、标定优先级排序以及在具有挑战性的地下场景中的可靠运行提供了面向部署的指导。为支持可复现评估和未来扩展,我们将发布完整的基准测试脚本和评估流程。
英文摘要
Visual-inertial odometry (VIO) is a core capability for autonomous operation in GPS-denied subterranean environments, yet its reliability can degrade sharply under sensor drift, calibration errors, and dynamic occlusion. Existing evaluations mainly emphasize nominal-condition accuracy, offering limited insight into when practical deployment failures occur. In this work, we present a failure-centric stress-test benchmark for VIO in underground environments using the CERBERUS dataset. We systematically evaluate four representative VIO systems spanning filtering-, optimization-, and learning-based paradigms under nine practical perturbation settings, including IMU bias and noise variation, camera intrinsic and extrinsic drift, and dynamic scene occlusion. Beyond conventional trajectory error, we analyze robustness limits through coverage ratio and failure thresholds, revealing breakdown behaviors that are not captured by nominal-condition performance alone. Our study shows distinct vulnerability patterns across VIO paradigms: some methods are more sensitive to inertial degradation, while others are more affected by geometric miscalibration or dynamic interference. These results provide deployment-oriented guidance for VIO selection, calibration prioritization, and reliable operation in challenging underground scenarios. To support reproducible evaluation and future extensions, we will release the full benchmark scripts and evaluation pipeline.
Comments8 pages, 5 figures