发表机构
University of California, Berkeley; Sharpa Robotics; The University of Hong Kong(加州大学伯克利分校; Sharpa Robotics; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出将生成视频与仿真落地结合的方法,通过HOI重建生成多样化参考,训练多对象跟踪器,在仿真中成功率提升超25个百分点,实现真实世界多样化灵巧操作。
AI 中文摘要
生成的手-物交互(HOI)视频为提出操作运动提供了一种可控的方式。基于仿真的HOI跟踪可以将此类运动学参考转化为可行的底层控制,但其可扩展性受限于缺乏可靠的参考运动。因此,我们将生成的视频与基于仿真的HOI落地相结合:在训练期间,生成的视频为学习多对象、多轨迹HOI跟踪器提供多样化的运动参考;在部署时,视频模型生成运动计划,由学习到的跟踪器执行。具体而言,我们提出了一种方法,通过最少的人工干预进行HOI重建,实现可扩展的参考生成,并成功将超过1,500个生成的视频落地到仿真中,在基于仿真的训练中,成功率比基线高出超过25个百分点。在真实世界的闭环实验中,它实现了多样化的抓取,包括功能性抓取、非抓握操作和抓取后对象姿态跟踪。视频和代码可在该https URL获取。
英文摘要
Generated hand-object interaction (HOI) videos provide a controllable way to propose manipulation motions. Simulation-based HOI tracking can translate such kinematic references into feasible low-level control, but its scalability is limited by the lack of reliable reference motions. We therefore combine generated videos with simulation-based HOI grounding: during training, generated videos provide diverse motion references for learning a multi-object, multi-trajectory HOI tracker, and at deployment, the video model produces motion plans that are executed by the learned tracker. In particular, we propose a method that enables scalable reference generation by HOI reconstruction with minimal manual intervention and successfully grounds more than 1,500 generated videos in simulation, achieving success rates over 25 percentage points higher than those of baselines during simulation-based training. In real-world closed-loop experiments, it achieves diverse grasps, including functional grasps, non-prehensile manipulation, and post-grasp object-pose tracking. Videos and code are available at https://boyuan-an.github.io/GALATEA/.
CommentsProject website: https://boyuan-an.github.io/GALATEA/