arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21302cs.AI

专家行为先验强化学习

Expert Behavior Prior Reinforcement Learning

  • School of Computer Science, Tongji University(同济大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia

AI总结:

研究针对行为先验强化学习依赖静态离线数据集的问题,提出专家行为先验算法,通过Q引导的条件变分自编码器生成专家策略先验,并结合专家策略引导和策略梯度校正模块,提升样本效率与收敛稳定性。

AI中文摘要:

行为先验强化学习(BPRL)是一种通过利用离线示范得出的策略先验来提高在线强化学习样本效率的范式。但现有大多数BPRL方法依赖静态离线数据集,存在数据多样性低和轨迹质量次优的问题,限制了策略先验的有效性。为此提出专家行为先验(EBP)算法,引入Q引导的条件变分自编码器直接从在线重放缓冲区生成专家策略先验,还提出专家策略引导机制和策略梯度校正模块。实验表明EBP显著优于现有算法,样本效率更高且收敛更稳定。

英文摘要:

Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations. However, most existing BPRL methods rely on static offline datasets, which often suffer from low data diversity and suboptimal trajectory quality. This reliance restricts the effectiveness of policy priors, hindering both policy exploitation and stability during online training. Consequently, agents are prone to inefficient exploration and unstable learning dynamics. To address these limitations, we deviate from existing offline pre-training methods and propose an Expert Behavior Prior (EBP) algorithm. Specifically, we introduce a Q-guided conditional variational autoencoder (Q-CVAE) that learns to generate expert policy priors directly from the online replay buffer. This enables the generation of high-value actions for guiding policy updates without relying on pre-collected expert trajectories. To further enhance policy exploitation, we propose an expert policy guidance (EPG) mechanism that selects expert actions from a generative support set, and we integrate a policy gradient correction (PGC) module to harmonize Q-guidance with expert supervision, promoting stable and consistent policy improvement. Extensive experiments conducted on robotic control (Gym, PyBullet) and industrial control (DMControl) benchmarks demonstrate that EBP significantly outperforms state-of-the-art online RL algorithms, achieving higher sample efficiency and more stable convergence.

补充信息

↑