arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于分层动作分块的离线强化学习

Offline RL with Hierarchical Action Chunking

Ahad Jawaid

arXiv 2607.20834首次发表:更新:

AI 中文总结

研究离线目标条件强化学习中扩展到长期任务的挑战,提出HiQC算法,结合高级潜在规划与低级动作分块,实现无偏k步价值备份,理论证明其价值误差界限更紧,实证表明在OGBench套件中性能最佳。

AI 中文摘要

离线目标条件强化学习有望从静态数据集中学习通用策略。然而,由于视界诅咒,将这些方法扩展到长期任务仍是挑战,价值估计误差会在引导式贝尔曼备份的长链中累积。现有分层方法通过将任务分解为子目标来缓解此问题,但常依赖存在近视执行和有偏价值估计的低级控制器。本文提出分层隐式Q分块(HiQC),一种结合高级潜在规划与低级动作分块的离线目标条件强化学习算法。通过对低级评论家基于时间扩展动作序列进行条件设定,HiQC实现无偏k步价值备份,在规划和执行层面压缩视界。理论上证明,与单独的标准分层或平面分块相比,这种双重分解在有界每次备份误差模型下导致更紧的价值误差界限。实证上,HiQC在OGBench套件的比较方法中实现最高总体性能,在长期导航任务如人形巨人任务上增益最大。

英文摘要

Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups. Existing hierarchical approaches mitigate this by decomposing tasks into subgoals, yet they often rely on low-level controllers that suffer from myopic execution and biased value estimates. In this work, we propose Hierarchical Implicit Q-Chunking (HiQC), an offline goal-conditioned RL algorithm that combines high-level latent planning with low-level action chunking. By conditioning the low-level critic on temporally extended action sequences, HiQC enables unbiased k-step value backups, compressing the horizon at both the planning and execution levels. We theoretically demonstrate that this dual decomposition results in a tighter bound on value error under a bounded per-backup error model compared to standard hierarchy or flat chunking alone. Empirically, HiQC achieves the highest aggregate performance among the compared methods on the OGBench suite, with its largest gains on long-horizon navigation tasks such as humanoid-giant.

CommentsRLC/RLJ 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑