arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FetchMan:基于模拟经验学习人形机器人视觉 locomotion-manipulation(移动操作)策略

FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences

Omar Rayyan, Zhi Li, Max Argus, Yuxin Jiang, Chang Yu, Chenfanfu Jiang, Yuchen Cui

arXiv 2608.17027首次发表:更新:

发表机构

University of California, Los Angeles; Allen Institute for AI; University of Washington(加州大学洛杉矶分校; 艾伦人工智能研究所; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 FetchMan,通过端到端 sim-to-real 流水线训练人形机器人视觉移动操作策略,在 FetchMan-Bench 评估中,其单物体抓取策略在 Unitree G1 上零样本部署成功率达73.3%,并扩展至多物体训练。

AI 中文摘要

能够泛化到新场景和新物体的视觉移动操作策略长期以来一直是机器人研究的目标。然而,如今数据密集型算法使得桌面操作的充足演示数据收集变得困难,对于还需要行走和平衡的人形机器人来说更是如此。像 locomotion(移动)领域常见的那样,从模拟数据中学习并将行为迁移到现实世界可以规避这一难题,因此我们将该方法应用于移动操作领域。在实践中我们发现,无论训练数据量多大,克隆合成演示都会导致性能上限较低。强化学习能够突破这一限制,通过在单一稀疏奖励下使用 Flow-GRPO 优化克隆策略,其性能优于合成行为克隆。这些阶段共同构成了我们的端到端 sim-to-real(模拟到现实)流水线,覆盖超过 150000 个场景,我们用它来训练 FetchMan。我们在自行发布的模拟基准 FetchMan-Bench 上对其进行评估,并将其零样本部署到现实世界的 Unitree G1 机器人上,其中我们的单物体抓取策略在未见场景中行走并抓取目标的成功率为 73.3%。最后,我们将该方法扩展到多物体训练,这是在该数据规模下实现移动操作通用策略的第一步。

英文摘要

Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robotics research. However, today's data-hungry algorithms make collecting sufficient demonstrations a struggle for tabletop manipulation, and even more so for humanoids that must also walk and balance. Learning from simulated data and transferring that behavior to the real world, as is commonly done in locomotion, sidesteps this struggle, so we replicate that recipe for loco-manipulation. In doing so, we find that cloning synthetic demonstrations results in a low performance ceiling no matter the amount of training data. Reinforcement learning breaks through it, and refining the cloned policy with Flow-GRPO on a single sparse reward yields performance that synthetic behavior cloning cannot match. Together, these stages form our end-to-end sim-to-real pipeline spanning more than 150,000 scenes, which we use to train FetchMan. We evaluate it on FetchMan-Bench, a simulation benchmark we release, and deploy it zero-shot on a real Unitree G1, where our single-object reach-and-pick policy walks to and grasps a target across unseen scenes at 73.3% success. Finally, we extend this recipe to multi-object training, a first step toward loco-manipulation generalist policies at this data scale.

CommentsProject website: https://orayyan.com/fetchman

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑