arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21735cs.LG

GEM-MPC:通过专家引导规划平衡探索与利用

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

Alvaro Serra-Gomez, Thomas Moerland

首次发表
浏览论文内容

中文总结 AI 辅助

GEM-MPC提出一种基于MPPI的强化学习方法,通过结合克隆规划器的策略与KL正则化探索策略,并引入门控先验蒸馏,在较低计算预算下平衡探索与利用,显著提升连续控制基准性能。

中文摘要 AI 辅助

在高维连续控制中,有效的探索仍然是强化学习的核心挑战。基于规划的方法通过将在线规划与学习到的策略和价值函数相结合来解决这一问题,但它们的组件在训练过程中可能变得不对齐:学习到的采样策略可能偏离规划器的行为,而随着模型和价值函数的演化,存储在回放缓冲区中的规划分布会变得过时。重新分析可以刷新这些目标,但代价是巨大的计算开销。我们提出了GEM-MPC,一种基于MPPI的强化学习方法,旨在改善规划与学习之间的交互。GEM-MPC使用MPPI将克隆规划器的策略与围绕其进行探索的KL正则化策略相结合,在规划内提供互补的利用和引导式探索。我们进一步引入了门控先验蒸馏(Gated Prior Distillation),它仅在存储的规划分布比当前先验提供更好目标时选择性地从中学习,从而在无需完全重新分析的情况下减少过时规划数据的影响。在连续控制基准测试中,GEM-MPC在较低的计算预算下始终优于现有的基于规划的基线方法。

英文摘要

Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves the interaction between planning and learning. GEM-MPC uses MPPI to combine a policy trained to clone the planner with a KL-regularized policy that explores around it, providing complementary exploitation and guided exploration within planning. We further introduce Gated Prior Distillation, which selectively learns from stored planning distributions only when they provide a better target than the current prior, reducing the impact of stale planning data without requiring full reanalysis. Across continuous-control benchmarks, GEM-MPC consistently outperforms existing planning-based baselines under lower computational budgets.

发表机构

  • Leiden University(莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑