arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24145cs.RO

MimicAgent:通过提示到轨迹生成实现四足技能

MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Lucky Kant Nayak, Narayanan Palghat Parameswaran, Neehar Peri, Deva Ramanan

AI总结:

MimicAgent通过提示到轨迹生成框架,利用编码智能体生成四足参考轨迹,训练示例引导的强化学习策略,实现动态四足技能,87%提示生成语义对齐轨迹。

AI中文摘要:

我们提出了MimicAgent,一个用于学习动态四足技能的提示到轨迹生成框架。尽管在训练四足策略时广泛使用奖励塑形,但导航由此产生的奖励景观是出了名的困难,需要花费数小时的“研究生下降”过程。Eureka尝试用LLM自动化奖励设计,但我们发现它难以在不同技能和形态之间泛化。我们的关键观察是,对于人类——以及相关的LLM——生成参考运动比塑形奖励函数要容易得多。我们的假设受到人类形机器人示例引导强化学习成功的启发,该成功利用大规模动作捕捉数据集作为训练运动策略的参考。与人类形机器人不同,四足机器人缺乏此类参考运动数据。为此,我们提出了MimicAgent,一个智能体框架,给定技能提示,通过编码智能体生成四足参考轨迹。这些粗略的参考轨迹随后用于训练示例引导的强化学习策略,这些策略可在仿真和现实世界中部署。值得注意的是,我们发现在我们的智能体框架内提示Claude Fable 5.1时,87%的提示能产生语义对齐的参考轨迹。

英文摘要:

We present MimicAgent, a prompt-to-trajectory generation framework for learning dynamic quadruped skills. Although reward shaping is extensively used when training quadruped policies, navigating the resulting reward landscape is notoriously difficult, requiring hours of "graduate student descent". Eureka attempts to automate reward design with LLMs, but we find that it struggles to generalize across diverse skills and morphologies. Our key observation is that it is far easier for a human - and by association, an LLM - to generate reference motions than to shape reward functions. Our hypothesis is motivated by the success of example-guided RL for humanoids, which exploits large-scale motion capture datasets as references for training locomotion policies. Unlike humanoids, quadrupeds lack such reference motion data. Towards this end, we propose MimicAgent, an agentic harness that, given a skill prompt, generates quadruped reference trajectories with coding agents. These coarse reference trajectories are then used to train example-guided RL policies that are deployable in simulation and in the real-world. Notably, we find that when prompting Claude Fable 5.1 within our agentic harness, 87% of prompts yield semantically aligned reference trajectories.

补充信息

↑