arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VOMMI:收集与利用便携式演示进行移动操作

VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile Manipulation

Yutian Zhang, Xingrui Xiong, Siyuan Ma, Yang Li, Jiawen Wen, Jiaqi Zhai, Liwen Yang, Ce Hao, Haozhen Chi, Yangkun Zhu, Yifan Zhu, Xiaowen Chu, Dong Wei, Qiaojun Yu, Dibo Hou

arXiv 2610.08220首次发表:更新:

发表机构

ZJU; Shanghai AI Lab; DeepRobotics; Yale; THU; ZGCA; HKUST(GZ); ZUST(浙江大学; 上海人工智能实验室; 深圳优必选科技股份有限公司; 耶鲁大学; 清华大学; 中国美术学院; 香港科技大学(广州); 浙江科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VOMMI框架,通过离线轨迹重建与在线视觉-运动条件化,利用便携式RGB演示进行移动操作策略后训练,降低速度误差并提升成功率。

AI 中文摘要

便携式移动操作演示有助于缓解具身智能的数据稀缺问题,但从RGB观测中获取可靠、低成本且无需机器人的运动监督仍然具有挑战性。现有方法通常依赖遥操作或配备额外传感硬件的专用设备,而直接使用估计的视觉里程计(VO)轨迹可能因累积漂移和不完美的运动监督而产生不一致。我们提出了视觉里程计条件移动操作接口(VOMMI),一个便携式演示收集与学习框架,通过离线轨迹重建和在线视觉-运动条件化,将便携式RGB演示与视觉-语言-动作(VLA)后训练连接起来。VOMMI同步身体和手部视图以捕获导航上下文和局部物体交互,无需人体-机器人运动学对应校准。R2-VO利用稀疏几何锚点细化离线演示轨迹,并在多个预测视界上生成因果局部运动令牌,用于在线策略条件化。一个动作组残差适配器仅将这些令牌纳入基础分支。实验为每个任务使用500条便携式轨迹,其中75条轨迹留出用于RGB-VO评估,200条机器人演示作为参考。我们的策略仅在后训练于便携式演示上,其基础速度误差比使用机器人收集演示训练的策略低18.2%,同时保持可比的末端执行器平移精度。离线重建使身体和手部流的绝对轨迹误差相对于每个流的最佳评估基线平均降低24.6%。完整系统在三个真实机器人任务上,相比OpenPI 0.5,平均成功率提高了8.3个百分点。

英文摘要

Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision. We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning. VOMMI synchronizes body and hand views to capture navigation context and local object interactions without requiring human-robot kinematic correspondence calibration. R2-VO refines offline demonstration trajectories using sparse geometric anchors and produces causal local-motion tokens over multiple prediction horizons for online policy conditioning. An action-group residual adapter incorporates these tokens only into the base branch. Experiments use a 500-trajectory portable for each task, with 75 trajectories held out for RGB-VO evaluation, and 200 robot demonstrations as references. Our policy, post-trained only on portable demonstrations, achieves 18.2% lower base-velocity error than a policy trained with robot-collected demonstrations, while maintaining comparable end-effector translation accuracy. Offline reconstruction reduces absolute trajectory errors for the body and hand streams by 24.6% on average relative to the best evaluated baseline for each stream. The complete system improves the mean success rate by 8.3 percentage points over OpenPI 0.5 across three real-robot tasks.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑