arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

离线强化学习中的策略提取解耦

Decoupling Policy Extraction for Offline Reinforcement Learning

Xuyao Lin, Yixiang Shan, Jinru Duan, Tao Yang, Xinyu Zhao, Runyu Lei, Yiming Zhao, Jiaxin Fan, Zongbao Feng, Peng Jia

arXiv 2608.20909首次发表:更新:

发表机构

Simplexity Robotics; Rensselaer Polytechnic Institute; Northeastern University(森普莱西蒂机器人公司; 伦斯勒理工学院; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对离线强化学习中演员与评论员耦合训练的问题,提出将策略改进与演员训练解耦的范式,通过演员生成动作候选、评论员选择,实验显示其性能优于行为克隆及联合训练的离线RL方法。

AI 中文摘要

离线强化学习(Offline RL)方法通常联合训练演员(actor)和评论员(critic),其中评论员用于引导演员选择更高价值的动作。这种耦合学习过程在在线强化学习中是合理的,因为改进后的演员会收集新数据,可进一步更新演员和评论员。然而,离线强化学习中的训练数据是固定的,演员侧的策略改进无法生成新数据来验证或修正评论员。此外,保留这种耦合范式会带来两个相关挑战:其一,演员更新可能会漂移到高价值但可能分布外(OOD)的动作,并放大评论员的高估问题;其二,保守价值估计或行为克隆正则化会在抑制分布外动作与在数据支持区域内选择高价值动作之间造成艰难的权衡。基于这一观察,我们重新审视传统离线强化学习范式,提出将策略改进与演员训练解耦的方法。具体而言,我们仅训练演员来建模行为分布,并在推理时通过用单独学习的评论员对多个演员生成的提议进行重新排序来执行策略改进,我们将这种范式称为解耦策略提取范式。在该范式下,演员提供行为支持的动作候选,而评论员在该候选集中执行基于价值的选择。大量实验表明,解耦策略提取范式的性能优于行为克隆和联合训练的离线强化学习方法,即使使用朴素的Q学习评论员也能保持有效。

英文摘要

Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled learning process is well motivated in online RL, where an improved actor collects new data that can further update the actor and the critic. However, training data remains fixed in offline RL, making actor-side policy improvement unable to generate new data to validate or correct the critic. Moreover, retaining this coupled paradigm leads to two related challenges. Firstly, actor updates can drift toward high-valued but potentially out-of-distribution (OOD) actions and amplify critic overestimation. Secondly, conservative value estimation or behavior-cloning regularization creates a difficult trade-off between suppressing OOD actions and selecting high-value actions within the data-supported region. Motivated by this observation, we revisit the conventional offline RL paradigm and propose decoupling policy improvement from actor training. Specifically, we train the actor solely to model the behavior distribution and perform policy improvement at inference time by reranking multiple actor-generated proposals with a separately learned critic. We refer to this paradigm as the decoupled policy extraction paradigm. Under such paradigm, the actor provides behavior-supported action candidates, while the critic performs value-based selection within this candidate set. Extensive experiments show that the decoupled policy extraction paradigm outperforms both behavior cloning and jointly learned offline RL methods, while remaining effective even with a naive Q-learning critic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑