arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40149cs.LG

角色自适应策略优化用于离线强化学习

Role-Adaptive Policy Optimization for Offline Reinforcement Learning

Seonvin Cho, Soohyun Choi, Songnam Hong

首次发表
浏览论文内容

中文总结 AI 辅助

针对离线强化学习中策略正则化角色耦合问题,提出角色自适应策略优化(RAPO),按价值学习与执行角色独立调整更新系数,在D4RL任务上超越TD3+BC和IQL基线。

中文摘要 AI 辅助

离线强化学习中的策略正则化在策略改进与对不确定价值估计的依赖之间寻求平衡。这种平衡在选择执行动作与为评论家自举提供动作之间可能有所不同,然而诸如TD3+BC之类的方法通过共享策略将这两种角色耦合在一起。我们提出了角色自适应策略优化(RAPO),该方法根据策略在价值学习和执行中的角色来调整策略更新系数。RAPO通过利用基础算法的演员目标对候选策略更新进行微分来学习这些系数。对于TD3+BC,RAPO将自举演员和执行演员分离,并独立调整它们的系数:自举目标惩罚由策略引起的目标价值变化,而执行目标评估局部策略改进的替代指标。对于IQL,其价值学习已经独立于执行演员,RAPO保留原始价值更新,仅调整优势加权策略提取中的逆温度参数。在D4RL运动任务和AntMaze任务上的实验表明,RAPO相对于两种基础算法均有改进,其中TD3+BC的改进更大,其RAPO实例在平均性能上优于基线。

英文摘要

Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.

发表机构

  • Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑