基于世界模型的潜在能量动作规划
Latent Energy Action Planning with World Models
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对潜在世界模型规划中动作序列与目标描述符不匹配的问题,提出LEAP方法,通过耦合终端潜在目标匹配与状态能量,结合拟牛顿求解器等,在四个控制领域将规划成功率提升17.3个百分点。
AI中文摘要:
潜在世界模型支持从高维观测中进行高效的模型预测控制,但优化单个学习到的潜在目标可能会倾向于产生这样的动作序列:其解码器预测的终端描述符与目标描述符不匹配。我们提出了潜在能量动作规划(Latent Energy Action Planning,LEAP),该方法将完整的动作 horizon 视为可微变量,并通过冻结的 LeWorldModel(LeWM)对其进行优化。LEAP 将终端潜在目标匹配与终端窗口状态能量耦合;低能量要求预测的终端潜在与目标潜在一致,且解码器预测的终端描述符与目标描述符一致。一个冻结的目标条件提议初始化搜索,拟牛顿求解器通过自回归展开优化动作,优化后投影则确保动作处于允许的范围内。在使用官方发布的 LeWM 检查点的四个控制领域中,完整的 LEAP 规划系统将 LeWM 结合交叉熵方法(LeWM+CEM)规划的平均成功率从 77.5% 提升至匹配协议下的 94.8%,实现了 17.3 个百分点的提升,同时保留了冻结的 LeWM 表示和预测器。
英文摘要:
Latent world models support efficient model predictive control from high-dimensional observations, yet optimizing a single learned latent objective can favor action sequences whose decoder-predicted terminal descriptor does not match the goal descriptor. We introduce Latent Energy Action Planning (LEAP), which treats the complete action horizon as a differentiable variable and optimizes it through a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy. Low energy requires the predicted terminal latent to agree with the goal latent and the decoder-predicted terminal descriptor to agree with the goal descriptor. A frozen goal-conditioned proposal initializes the search, a quasi-Newton solver refines actions through the autoregressive rollout, and post-optimization projection enforces the admissible action range. Across four control domains using the officially released LeWM checkpoints, the complete LEAP planning system raises mean success from 77.5% for LeWM planned with the cross-entropy method (LeWM+CEM) to 94.8% under a matched protocol, a 17.3-percentage-point improvement, while retaining the frozen LeWM representation and predictor.