发表机构
Chongqing University; Northwestern Polytechnical University; The University of Tokyo; Jiaxing University; Sichuan University; Beijing Institute of Technology; Shanghai Jiao Tong University(重庆大学; 西北工业大学; 东京大学; 嘉兴大学; 四川大学; 北京理工大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对UMI式机器人示教中末端执行器定位难题,提出MILD数据集与AprilVINS方法,结合鱼眼视觉惯性估计与AprilTag几何,实现毫米级定位精度,并提供诊断性基准框架。
AI 中文摘要
机器人示教学习需要在近距离操作和相机遮挡期间对末端执行器进行精确且时间完整的定位。现有的SLAM基准测试强调导航运动,而操作数据集则优先考虑策略学习而非定位评估。我们引入了MILD,一个包含真实世界和仿真序列的操作接口定位数据集。真实世界子集提供了来自Insta360 X5和Insight9的86个传感器序列,涵盖15个重复的桌面任务、标定资产以及每次执行的机器人末端执行器参考轨迹。仿真子集MILD-Sim在Isaac Sim中扩展了任务覆盖范围,用于受控的操作重放研究。在仪器化的真实世界记录上对视觉惯性系统和辅助标记系统进行基准测试,揭示了即使在相同名义任务下,TCP相对轨迹误差和时间覆盖度也存在巨大差异。为了支持无需预先测量的标记地图的标记增强示教工作空间,我们提出了AprilVINS,它结合了鱼眼视觉惯性估计与序列局部AprilTag几何,并将先验接纳与联合优化状态的受控导出分开。在Insta360 AprilTag4记录上,使用统一协议和序列特定配置的AprilVINS(full)达到了毫米级的SE(3)对齐TCP相对APE RMSE,具有高时间完成度,并且在各自协议下比测试路线报告的误差更低,而无需标签因子的鱼眼VIO仍保持在厘米级。消融研究将精度与可导出性分开,MILD-Sim重放研究为解释这些误差幅度提供了任务特定的容差参考。总之,MILD和AprilVINS为UMI式示教收集提供了一个诊断性基准测试框架。代码、数据集和评估清单将在录用后发布。
英文摘要
Robot demonstration learning requires accurate and temporally complete end-effector localization during close-range manipulation and camera occlusion. Existing SLAM benchmarks emphasize navigation motions, whereas manipulation datasets prioritize policy learning over localization evaluation. We introduce MILD, a Manipulation-Interface Localization Dataset with real-world and simulation sequences. The real-world subset provides 86 sensor sequences from Insta360 X5 and Insight9 across 15 repeated tabletop tasks, calibration assets, and a per-execution robot end-effector reference trajectory. The simulation subset, MILD-Sim, extends task coverage in Isaac Sim for controlled manipulation-replay studies. Benchmarking visual-inertial and fiducial-aided systems on instrumented real-world recordings reveals large differences in both TCP-relative trajectory error and temporal coverage, even under the same nominal task. To support marker-augmented teaching workspaces without a pre-surveyed fiducial map, we present AprilVINS, which combines fisheye visual-inertial estimation with sequence-local AprilTag geometry and separates prior admission from guarded export of the jointly optimized state. On Insta360 AprilTag4 recordings, AprilVINS(full) under a unified protocol with sequence-specific profiles reaches millimeter-level SE(3)-aligned TCP-relative APE RMSE with high time completion and lower reported error than the tested routes under their respective protocols, whereas fisheye VIO without tag factors remains at centimeter scale. Ablations separate accuracy from exportability, and a MILD-Sim replay study provides task-specific tolerance references for interpreting those error magnitudes. Together, MILD and AprilVINS provide a diagnostic benchmarking framework for UMI-style demonstration collection. Code, datasets, and evaluation manifests will be released upon acceptance.
Comments9 pages, 8 figures, 5 tables. Project page: https://mild-web.github.io