arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:用于潜在世界模型规划的子目标条件动作生成

SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning

Letian Cheng, Qi Zhang, Yisen Wang

arXiv 2607.17973首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究潜在世界模型规划中提议质量受限问题,提出先验条件规划器,利用目标条件生成器预测潜在子目标,结合不同时长子目标作先验,实验证明该方法显著提升长期规划性能。

AI 中文摘要

潜在世界模型已成为一种强大的规划范式,通过学习动作条件预测动力学并将其用作内部模拟器来想象和评估候选动作序列。然而,随着规划范围的扩大,性能越来越受到提议质量的限制。本文介绍了一种先验条件规划器,用结构化指导取代随机提议初始化。在每个规划阶段,目标条件生成器预测指定持续时间内的下一个可达潜在子目标,用于条件生成候选动作序列。为跨时间尺度捕获语义信息,使用不同持续时间的子目标作为先验。实验表明,耦合潜在子目标分解与先验条件动作生成可显著改善长期规划,同时保持强大的短期性能。

英文摘要

Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences. However, as the planning horizon grows, performance becomes increasingly constrained by proposal quality: a fixed candidate budget must search an exponentially larger action space, making it difficult to expose the world model to high-quality candidate futures for evaluation. In this paper, we introduce SAGE, a prior-conditioned planner that replaces random proposal initialization with structured guidance. At each planning stage, a goal-conditioned generator predicts the next intermediate latent subgoal for a specified duration, which is then used to condition the generation of candidate action sequences. To capture semantic information across temporal scales, we use subgoals of varying durations as priors, balancing fine-grained local control with higher-level long-horizon progress. Then the frozen world model evaluates these proposals against the same subgoal and guides their refinement before execution. Experiments on PushT and OGBench Cube show that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance. To be specific, when the target offset is $150$, it raises PushT success from $4.7\%$ to $64.7\%$ and OGBench Cube success from $20.7\%$ to $67.3\%$. We further extend latent world-model planning to LIBERO, where SAGE improves full-episode success from $0\%$ with the vanilla LeWM planner to $48.7\%$ on Scene2 and $58\%$ on Caddy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑