发表机构
University of California, Berkeley; ShanghaiTech University; Google DeepMind; Nomagic(加州大学伯克利分校; 上海科技大学; 谷歌DeepMind; Nomagic)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HuGo利用LLM从任务描述生成可执行的高层策略代码,结合冻结的低层全身策略,通过数值轨迹和视频帧反馈细化,无需任务特定奖励或演示,在模拟和硬件上显著优于强化学习基线。
AI 中文摘要
为了使类人机器人在日常环境中发挥作用,它们必须执行一系列将移动与操作相结合的任务。现有方法通常通过奖励工程或演示来获取移动操作策略,然后进行特定任务的训练,这使得扩展到新任务的成本高昂。在这项工作中,我们提出了一种用于人形机器人移动操作的分层方法,消除了这些每任务的需求。HuGo,即人形机器人策略代码生成,使用大型语言模型(LLM)根据任务描述在冻结的低层全身策略之上生成可执行的、闭环的高层策略代码。给定任务、观测和指令规格,LLM在代码中构建任务逻辑。HuGo随后利用数值轨迹和选定的视频帧从其展开中细化策略,以产生反馈和有针对性的代码更新。在五个模拟任务中,使用两种不同的低层策略,HuGo显著优于高层强化学习基线,并接近基于演示的基线的性能。我们在没有特定任务奖励设计或演示收集的情况下实现了这一性能水平。我们进一步展示了模拟生成策略到硬件的零样本迁移,并表明将相同的细化循环应用于真实世界展开可以进一步提高迁移性能,而无需专家演示或策略重新训练。项目网站位于此https URL。
英文摘要
For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/