arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36896cs.AI

HorizonFlow:离线目标条件强化学习的变长规划

HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

JunHyeok Oh, Zian Jang, Byung-Jun Lee

首次发表
浏览论文内容

中文总结 AI 辅助

HorizonFlow提出将规划长度作为生成输出的分层规划器,通过流匹配与插入式生成联合优化长度与内容,在多个基准上取得最优平均性能。

中文摘要 AI 辅助

生成式规划的最新进展使得轨迹修补成为离线目标条件强化学习的一种有前景的方法。然而,这些方法通常在生成规划内容之前指定规划范围,尽管合适的范围取决于路线本身。过短的范围可能迫使不可行的转换,而过长的范围则可能引入冗余运动。我们提出了HorizonFlow,一种分层规划器,它将规划长度视为生成的输出而非预设输入。其子目标路线规划器通过一系列潜在子目标引导其动作前缀控制器。两个组件都结合了基于插入的生成与流匹配,以联合生成连续的规划内容和长度,利用部分生成的规划来指导标记插入。HorizonFlow重用由此产生的长度信息来选择候选方案,并引导生成朝向更短的规划,而无需单独学习价值模型。在Maze2D、Multi2D和OGBench导航及视觉操作基准上,HorizonFlow在比较方法中实现了最高的平均性能。

英文摘要

Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.

发表机构

  • Korea University(高丽大学)
  • Gauss Labs Inc.(高斯实验室公司)

机构由 AI 辅助整理,请以论文原文为准。

↑