arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散子目标规划用于长视界离线目标条件强化学习

Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang

arXiv 2609.34575首次发表:更新:

发表机构

China University of Mining and Technology; South China University of Technology(中国矿业大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视界离线目标条件强化学习中价值函数不稳定问题,提出基于扩散的子目标生成框架DSP,通过无分类器引导生成可达且目标导向的子目标,在导航和操作任务上优于现有方法。

AI 中文摘要

离线目标条件强化学习(GCRL)从无奖励数据中学习目标导向策略,但在长视界任务中,由于稀疏奖励和折扣,目标条件价值函数往往提供不稳定的指导。分层方法通过子目标分解部分缓解了这一问题;然而,高层决策仍依赖于对噪声敏感的价值估计,导致在复杂环境中行为不稳定。我们通过提出扩散子目标规划(DSP)来解决这一局限,这是一种基于扩散的高层子目标生成框架。DSP将高层规划视为对目标条件子目标的引导式生成推理,并同时学习条件流和无条件流,从而在推理时通过无分类器引导引入目标导向偏差。通过从高层规划中移除显式的基于价值的引导,DSP通过生成模型产生可达且目标导向的子目标,同时保留分层执行。在离线GCRL基准上的实验表明,DSP在多种导航和操作任务上优于先前方法,在需要多步子目标规划的迷宫环境中表现尤为突出。

英文摘要

Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.

Comments31 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑