HiFi-UMI:仅从高保真UMI数据中学习可部署的操作策略
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Simple AI(简单人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究如何从高保真UMI数据学习可部署操作策略,提出HiFi-UMI系统,其无需外部跟踪达3毫米精度。利用该系统数据实现零机器人训练后策略,预训练还能降误差提成功率,且开源了相关高保真资源。
AI中文摘要:
学习可部署的操作策略受到高保真且可扩展数据稀缺的限制。真实机器人遥操作准确但扩展成本高;无机器人UMI捕获易于扩展,当前主要用于预训练,训练后添加少量真实机器人“锚点”。我们探讨提高无机器人UMI数据保真度能否去除该锚点。提出HiFi-UMI,一个针对轨迹精度、夹爪间相对姿态、同步和视野共同设计的便携式UMI数据生成系统。它无需外部跟踪基础设施就能达到3毫米的工作空间局部末端执行器精度。利用该数据集,我们展示了零机器人训练后策略:仅在HiFi-UMI演示上训练后的策略可直接部署在真实机器人上,并在跨越视觉-语言-动作和世界-动作-模型家族的三个主干上与领域内遥操作匹配。在预训练4000小时后,十个未见任务的动作误差降低41%,在StarVLA-QwenPI上真实机器人成功率进一步提高18.1个百分点。我们开源HiFi-UMI-2K,作为机器人学习社区的大规模、高保真资源。
英文摘要:
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.