发表机构
School of Chemistry, Chemical Engineering and Biotechnology, Nanyang Technological University(南洋理工大学化学、化学工程与生物技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对离线强化学习中分布偏移导致的Q值估计偏差问题,提出带行为优势修正的扩散策略(DPBAC)算法,结合BAC-PE与扩散模型,在D4RL任务上实现优于SOTA的性能。
AI 中文摘要
在离线强化学习(RL)中,行为数据与学习到的策略之间的分布偏移会导致Q值估计错误,从而误导策略优化方向。为解决该问题,我们提出了行为优势修正策略评估(BAC-PE)方法,利用行为策略的Q函数修正学习到的策略的Q函数,以此缓解悲观保守性和高估偏差。此外,我们从理论上分析了BAC-PE的收敛性,并推导了学习到的Q函数与真实Q函数之间差异的上界。为缓解分布偏移,本研究采用扩散模型同时表示行为策略和学习到的策略,执行分布匹配以实现准确的策略正则化。此外,我们将Q值引导纳入训练过程,以实现有效的策略改进。通过将BAC-PE与扩散策略建模相结合,我们提出了带行为优势修正的扩散策略(DPBAC)算法。与现有离线方法相比,DPBAC展现出更强的策略表示能力,且能有效缓解Q值估计中的偏差。在D4RL任务的多个领域上的实验结果表明,DPBAC实现了更优的性能,相较于最先进(SOTA)算法具有显著优势。
英文摘要
In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization. To address this issue, we develop a behavioral advantage corrected policy evaluation (BAC-PE) approach, which utilizes the \emph{Q}-function of the behavior policy to correct the learned policy's \emph{Q}-function, thus mitigating pessimistic conservatism and overestimation bias. Furthermore, the convergence of BAC-PE is analyzed theoretically, and an upper bound on the difference between the learned \emph{Q}-function and the true \emph{Q}-function is derived. To alleviate distribution shift, this work employs diffusion models to represent both the behavior policy and the learned policy, performing distribution matching for accurate policy regularization. Additionally, \emph{Q}-value guidance is incorporated into the training process to achieve effective policy improvement. By combining BAC-PE with diffusion policy modeling, we propose the diffusion policy with behavioral advantage correction (DPBAC) algorithm. Compared to existing offline methods, DPBAC demonstrates stronger policy representation capabilities and effectively mitigates the bias in \emph{Q}-value estimation. Experimental results on multiple domains of D4RL tasks show that DPBAC achieves superior performance, with notable advantages over state-of-the-art (SOTA) algorithms.