arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LePlanner:一种用于世界模型的迭代摊销控制器

LePlanner: An Iterative Amortized Controller For World Models

Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M

arXiv 2609.13845首次发表:更新:

发表机构

Indian Institute of Technology Roorkee(印度理工学院鲁尔基分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LePlanner提出一种摊销迭代控制器,利用冻结世界模型预测器构建并细化潜在动作序列,以到达保持目标解决水平重置拖延,在多项任务中匹配或超越搜索规划器且计算成本大幅降低。

AI 中文摘要

使用联合嵌入预测架构训练的世界模型可以从物理交互中学习紧凑、结构化的潜在表示,然而在这些潜在空间中进行规划通常依赖于两种代价高昂的方法之一。基于搜索的规划器(如CEM、MPPI和iCEM)通过多次预测器展开来优化动作序列,以高单次决策计算量和延迟为代价实现了强劲的性能。基于策略的方法将推理摊销为单次前向传播,但在接触丰富的任务中,当演示分布是多模态时,其性能可能会下降。我们提出了LePlanner,一种摊销的迭代控制器,它通过冻结的世界模型预测器学习构建和细化潜在动作序列。LePlanner使用到达并保持目标进行训练,该目标鼓励控制器在最早可行的时间范围内到达目标并保持在该处。这解决了水平重置拖延问题,这是一种反复的退缩水平重规划不断推迟目标到达的失败模式。额外的动作高斯损失使生成的动作保持在离线数据集的支持范围内。在导航、接触丰富的操作和连续控制环境中,LePlanner匹配或超过了基于搜索的规划器,同时所需的预测器评估次数少一个数量级,每次决策的墙钟时间低3-49倍。它在PushT上达到98%的成功率,在Reacher上达到100%,在TwoRooms上达到100%,在OGBench Cube任务上达到92%。这些结果表明,通过在线搜索发现的许多结构可以摊销到一个轻量级的迭代策略中,从而实现无需在线优化的快速、水平感知的非线性物理控制。

英文摘要

World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet planning in these latent spaces typically relies on one of two costly approaches. Search-based planners such as CEM, MPPI, and iCEM optimize action sequences through many predictor rollouts, achieving strong performance at the cost of high per-decision compute and latency. Policy-based methods amortize inference into a single forward pass but can degrade on contact-rich tasks where the demonstration distribution is multimodal. We propose LePlanner, an amortized iterative controller that learns to construct and refine latent action sequences through a frozen world-model predictor. LePlanner is trained with an arrival-and-hold objective that encourages the controller to reach the goal at the earliest feasible horizon and remain there. This addresses horizon-reset procrastination, a failure mode in which repeated receding-horizon replanning continually postpones goal arrival. An additional action-Gaussian loss keeps generated actions near the support of the offline dataset. Across navigation, contact-rich manipulation, and continuous-control environments, LePlanner matches or exceeds search-based planners while requiring an order of magnitude fewer predictor evaluations and 3-49x lower wall-clock time per decision. It achieves success rates of 98% on PushT, 100% on Reacher, 100% on TwoRooms, and 92% on the OGBench Cube task. These results show that much of the structure discovered through online search can be amortized into a lightweight iterative policy, enabling fast, horizon-aware, nonlinear physical control without online optimization.

Comments23 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑