arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向长 horizon 离线目标条件强化学习的递归价值学习

Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

Hyeonseong Jeon, Youngwoon Lee

arXiv 2609.02237首次发表:更新:

发表机构

Yonsei University; Seoul National University(延世大学; 首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对长 horizon 离线目标条件强化学习的难点,提出 DCRL 方法,通过平衡二叉树分解轨迹、递归训练价值,降低误差累积,在 OGBench 长 horizon 任务上取得优于基线的性能。

AI 中文摘要

将离线目标条件强化学习(GCRL)扩展到长 horizon 任务颇具难度,原因在于:其一,长程价值学习依赖的短程估计可能仍不准确;其二,基于最大值的价值备份会因反复传播而放大高估误差。本文提出 DCRL(分治强化学习,Divide-and-Conquer RL),该方法将每条轨迹段递归分解为平衡二叉树,从叶节点向根节点训练价值,因此每个父节点仅在其子节点完成更新后才进行更新,采用观测路径的精确分解,而非在噪声备选方案中选择。由于该目标学习的是演示路径上的价值(这些路径未必是最优的),DCRL 会在多条轨迹间联合传播价值以发现更短路径。得益于平衡二叉树,DCRL 将最坏情况的自举深度从线性降低至对数级,这种更短的依赖结构在实验中对应着慢得多的误差累积。在多种目标到达任务中,DCRL 的性能显著优于现有扁平离线 GCRL 方法;在 OGBench 的五个最具挑战性的长 horizon 任务上,它将现有最佳平均得分从 55 提升至 64,超越了所有扁平及分层基线方法。

英文摘要

Scaling offline goal-conditioned reinforcement learning (GCRL) to long-horizon tasks is difficult because (1) long-range value learning depends on shorter-range estimates that may still be inaccurate, and (2) max-based value backups can amplify overestimation through repeated propagation. We propose DCRL (Divide-and-Conquer RL), which recursively decomposes each trajectory segment into a balanced binary tree and trains the values from leaves to root. Each parent is therefore updated only after its children, using an exact factorization of the observed route rather than selecting among noisy alternatives. Since this objective learns values along demonstrated routes that are not necessarily optimal, DCRL jointly propagates values across trajectories to discover shorter routes. Thanks to the balanced binary tree, DCRL reduces worst-case bootstrap depth from linear to logarithmic, and this shorter dependency structure empirically corresponds to much slower error accumulation. Across diverse goal-reaching tasks, DCRL substantially outperforms prior flat offline GCRL methods, and on the five most challenging long-horizon OGBench tasks, it improves the best prior average score from 55 to 64, surpassing all flat and hierarchical baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑