发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对离线强化学习的策略更新难题,提出多步近端策略改进(MPI)方法,可在可控范围内提升策略性能,在 D4RL 基准上实现了对多种强离线基线的提升。
AI 中文摘要
离线强化学习(RL)必须协调两个相互竞争的要求:策略更新应保持在数据集支持的动作附近,以保持价值估计的可靠性,然而有意义的收益往往需要超出行为分布。我们通过将策略建模为具有选定度量几何的概率流形,开发了一种离线演员更新的几何视图。在此视角下,广泛类别的离线演员目标可被解释为单一近端策略改进步骤(SPI),即由评论家定义的能量诱导的流形梯度流的隐式离散化。基于这一见解,我们提出了多步近端策略改进(MPI),一种插件式细化机制,用于组合连续的重新中心化近端步骤。MPI 支持在数据集支持范围外进行可控策略改进,同时在每次细化时保留近端控制。该框架适配多种策略几何,并为确定性策略和对角高斯策略提供实用实例化。在 D4RL 基准上的实验表明,少量 MPI 细化可在多项任务上提升 TD3+BC、ReBRAC 和 IQL 等强离线基线。聚焦诊断进一步区分了重新中心化细化与固定目标更新调度,并表征了评论家误差下的局限性。
英文摘要
Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement. We study how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy. We parameterize the procedure by a nominal total horizon $T$ and $K$ stages with local horizon $T/K$, distinguishing subdivision at a fixed total horizon from additional refinement at a common local horizon. Our analysis shows that sequential re-centering can reach endpoints unavailable to any single proximal step and characterizes how subdivision reduces the leading local discretization error of ideal updates under a fixed critic. We consider TD3+BC and IQL-based policy extraction to examine how improvement composition interacts with actor objectives and policy geometry. TD3+BC experiments on D4RL locomotion suggest that subdivision can broaden the range of useful total horizons, while adding refinement stages at a fixed small local horizon can improve return. The results identify improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and additional policy extraction.
CommentsPreprint; 22 pages. Major revision with a new title, revised analysis of horizon subdivision and re-centering, and expanded experiments and controls