arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越策略对齐:利用学习的世界模型闭环机器人控制中的规划-学习回路

Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models

Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias

arXiv 2609.39751首次发表:更新:

发表机构

Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PL-MPC,通过混合TD目标、分歧感知终端估计和返回加权蒸馏,改进TD-MPC的规划-学习闭环,在HumanoidBench上显著提升性能,并实现零样本仿真到现实迁移。

AI 中文摘要

利用学习的世界模型进行规划,将在线轨迹优化与学习到的价值函数和策略函数相结合,用于高维控制。由于规划器决定了用于学习的经验,而学习到的评论家和行动者反过来又对未来的计划进行评分和提出,因此规划与学习形成了一个闭环反馈回路。TD-MPC 是这一设计的典型实例。最近的政策约束变体通过将学习到的策略与规划器行为对齐来加强回路的一部分。我们引入了 PL-MPC(规划-学习 MPC),它额外修改了评论家监督和规划器终端价值估计。混合多步 TD 目标使评论家更新在引导之前暴露于更多已实现的奖励;基于分歧的终端估计减少了 MPPI 规划期间不确定评论家价值的影响;以及返回加权行动者蒸馏强调了来自高回报回合的规划器执行动作。世界模型架构和 MPPI 优化器在其他方面保持不变。在 HumanoidBench 上,最大的增益出现在 balance-hard 任务上,总平均回报(TAR)从 98±18 增加到 387±255,以及 hurdle 任务,从 199±13 增加到 466±200;更广泛基准上的性能仍然依赖于任务,而 PL-MPC 在 DMControl 上保持竞争力。受控消融实验显示了两个任务上不同组件之间的交互。我们进一步在 7 自由度 KUKA IIWA14 上展示了扳手螺母对齐的零样本仿真到现实迁移,在训练对象尺寸和两个未见尺寸上获得了比 TD-M(PC)^2 更高的观察成功率。代码和数据将在以下网址提供:此 https URL。

英文摘要

Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed feedback loop. TD-MPC is a prominent instance of this design. Recent policy-constrained variants strengthen one part of the loop by aligning the learned policy with planner behavior. We introduce PL-MPC (Planning-Learning MPC), which additionally modifies critic supervision and planner terminal-value estimation. Hybrid multi-step TD targets expose critic updates to more realized rewards before bootstrapping; disagreement-aware terminal estimates reduce the influence of uncertain critic values during MPPI planning; and return-weighted actor distillation emphasizes planner-executed actions from high-return episodes. The world-model architecture and MPPI optimizer are otherwise unchanged. On HumanoidBench, the largest gains occur on \texttt{balance-hard}, where Total Average Return (TAR) increases from $98\pm18$ to $387\pm255$, and \texttt{hurdle}, from $199\pm13$ to $466\pm200$; performance across the broader benchmark remains task dependent, and PL-MPC remains competitive on DMControl. Controlled ablations show different component interactions across the two tasks. We further demonstrate zero-shot sim-to-real transfer on wrench-nut alignment with a 7-DoF KUKA IIWA14, obtaining higher observed success than TD-M(PC)^2 on the training object size and two unseen sizes. Code and data will be available at: https://pl-mpc-humanoid.github.io.

Comments8 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑