Metro-WM:具有可实现子目标的长时域潜在规划
Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals
浏览论文内容
中文总结 AI 辅助
Metro-WM通过从经验中检索真实状态作为子目标,构建图进行长时域规划,解决了潜在子目标不可实现的问题,在成功率、速度和计算效率上显著优于现有方法。
中文摘要 AI 辅助
基于联合嵌入预测架构(JEPA)的模型预测控制提供了一种强大的零样本目标到达规划器,但其仅在短规划时域内有效。分层扩展试图通过学习一个宏观规划器来预测中间潜在子目标以引导微观规划器,从而弥合这一差距。在本工作中,我们证明了无约束的潜在子目标预测存在根本性缺陷。一项严格的评估表明,一个领先的最先进的宏观规划器经常发出物理上不可实现的子目标。为解决此问题,我们引入了Metro-WM,一种分层框架,它通过从先前经验中检索真实状态而非生成无根据的潜在向量来发出子目标。具体而言,Metro-WM构建了一个图,其顶点是来自离线专家演示或随机动作轨迹的观测帧,允许来自不同情节的帧被连接并拼接成通往目标的路线。对整个图进行规划还使系统对执行错误具有高度鲁棒性:如果微观规划器偏离路线,Metro-WM会立即从当前状态找到新的最优路径。我们的实验表明,Metro-WM在长时域成功率上优于次优分层方法高达37.33个百分点,同时速度提升高达10.9倍,所需离线计算量减少13-56倍,且需要更少的超参数调优。额外分析显示,Metro-WM找到的路径比离线演示更短,优于依赖查询自身演示的预言机,并在极其稀疏的数据集条件下保持稳健性能。
英文摘要
Model-predictive control with Joint-Embedding Predictive Architectures (JEPAs) provides a strong zero-shot goal-reaching planner, but it is only effective over short planning horizons. Hierarchical extensions attempt to bridge this gap by learning a macro planner to predict intermediate latent sub-goals to guide the micro planner. In this work, we demonstrate that unconstrained latent sub-goal prediction is fundamentally flawed. A rigorous evaluation reveals that a leading state-of-the-art macro planner routinely emits physically unrealisable sub-goals. To resolve this, we introduce Metro-WM, a hierarchical framework that issues sub-goals by retrieving genuine states from prior experience rather than generating ungrounded latent vectors. Specifically, Metro-WM constructs a graph whose vertices are observed frames from offline expert demonstrations or random-action trajectories, allowing frames from different episodes to be connected and stitched into routes to the goal. Planning over the full graph also makes the system highly robust to execution errors: if the micro planner drifts off course, Metro-WM instantly finds a new optimal path from the current state. Our experiments show that Metro-WM achieves superior long-horizon success rates of up to 37.33 percentage points over the next best hierarchical approach while being up to 10.9 times faster, requiring both 13-56 times less offline compute and fewer tuned hyperparameters. Additional analysis reveals that Metro-WM finds shorter paths than the offline demonstrations, outperforms an oracle relying on the query's own demonstration, and maintains robust performance under extremely sparse dataset conditions.