arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教机器狗新技能:结合强化学习与模仿学习及对抗性任务选择实现四足机器人多样技能

Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

Lemon Foxmere, Anthony Furman, Yizheng Du, Oliver Chang, Leilani Gilpin, Steve McGuire

arXiv 2610.10601首次发表:更新:

发表机构

University of California at Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出结合强化学习、模仿学习与对抗性任务选择的三阶段方法,训练单一策略完成22项四足任务,在Unitree B1机器人上验证了优于基线的性能与现实鲁棒性。

AI 中文摘要

强化学习(RL)已使腿式机器人在单任务场景下完成一系列技能,但农业机器人、太空探索等应用需要 locomotion(移动)、挖掘、近距离勘测等多样技能。由于多任务学习中样本效率低、任务间梯度冲突等挑战,训练端到端策略解决该问题仍很困难。本文提出一种三阶段方法,训练单一策略完成行走、挖掘、跳跃等不同任务,并将其组合成爬行等新行为。首先,在窄定义任务上用 RL 训练多个教师策略;然后,通过两个额外阶段,在对抗性任务选择过程下,采用结合 RL 与模仿学习(IL)目标的多教师蒸馏设置训练学生策略,该过程将训练聚焦于表现最差的任务。通过此方法,用 8 个教师训练出执行 22 项任务的学生策略。评估显示,与 PPO 和蒸馏后微调基线相比,本文方法能保持运动质量、更准确跟踪指令,部分情况下可泛化至未明确训练的新任务。最后,将所得策略部署在 Unitree B1 四足机器人上,验证了其现实世界鲁棒性。视频链接:this https URL

英文摘要

Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional stages train a student policy with a multi-teacher distillation setup that uses a combined RL and Imitation Learning (IL) objective under an adversarial task selection process that focuses training on the worst-performing task. With this method, we train a student policy that performs 22 tasks using 8 teachers. Evaluations show our method preserves motion quality and tracks commands more accurately than PPO and distill-then-finetune baselines, and in some cases generalizes to new tasks without explicit training. Finally, we demonstrate real-world robustness by deploying the resulting policy on a Unitree B1 quadruped. Video: https://youtu.be/V9yX04EBcFA

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑