发表机构
Yale University; University of Sydney(耶鲁大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出采样引导策略搜索(SGPS),结合基于采样的模型预测控制与一阶策略优化,加速视觉策略学习,并在模拟及真实Go2机器人上实现零样本迁移。
AI 中文摘要
学习用于移动和操作任务的视觉策略需要与环境进行接触协调,并且可能产生大量的计算和GPU内存成本。一阶策略梯度(FoPG)通过可微仿真降低了训练成本,但局部优化可能收敛到非预期的接触模式。为解决这一不足,我们提出了采样引导策略搜索(SGPS),该方法将基于采样的模型预测控制所进行的周期性动作目标细化与一阶策略优化相结合。行为克隆利用采样动作初始化策略;随后,训练在扰动初始状态和随机动力学条件下,交替进行基于采样的细化与短视界FoPG更新。对于视觉策略训练,我们采用一种解耦的FoPG公式,该公式将渲染排除在计算图之外,从而无需状态策略教师即可直接从深度观测中学习。在单个GPU上,SGPS在模拟的Unitree Go2和G1机器人上学习了移动、障碍穿越、推箱子和双臂搬运的策略。我们的实验进一步表明,细化不仅优于仅初始化和仅跟踪,还能改善策略学习。对于硬件部署,蒸馏后的策略可零样本迁移到真实的Go2机器人上,并利用机载深度传感器自主执行小跑、爬行、跨栏以及在这些行为之间切换。
英文摘要
Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact patterns. To address this shortfall, we propose Sampling-Guided Policy Search (SGPS), which couples recurring action-target refinement by sampling-based model-predictive control with first-order policy optimization. Behavior cloning initializes the policy from sampled actions; training then alternates sampling-based refinement with short-horizon FoPG updates under perturbed initial states and randomized dynamics. For visual policy training, we use a decoupled FoPG formulation that excludes rendering from the computation graph, enabling direct learning from depth observations without a state-policy teacher. On a single GPU, SGPS learns policies for locomotion, obstacle traversal, crate pushing, and bimanual carrying on simulated Unitree Go2 and G1 robots. Our experiments further show that refinement improves policy learning beyond initialization and tracking alone. For hardware deployment, the distilled policy transfers zero-shot to a real Go2 and uses onboard depth to autonomously trot, crawl, clear hurdles, and switch between these behaviors.
Comments8 pages, 6 figures