通过双向行为先验蒸馏改进在线强化学习
Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation
查看机构详情
- Tongji University(同济大学)
- Shanghai University of Engineering Science(上海工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出双向行为先验蒸馏(B2PD)算法,利用动作价值先验引导CVAE生成高价值行为支持集并蒸馏到智能体,以缓解在线强化学习中的评论家误差放大,显著提升样本效率和策略稳定性。
中文摘要 AI 辅助
在线强化学习(RL)算法经常表现出较差的样本效率和动荡的学习动态,这源于系统性的评论家估计误差,而贪婪的策略更新会加剧这些误差。现有的行为先验强化学习方法试图通过依赖离线预训练从固定数据集中学习行为模型,并使用策略先验来约束在线策略更新,以缓解这一问题。然而,离线数据集的有限质量常常阻碍提供能够有效指导策略更新的高价值策略。专家轨迹的缺失严重损害了在线策略学习,导致样本效率低下和性能次优。为应对这些挑战,我们偏离传统的行为先验方法,提出了一种双向行为先验蒸馏(B2PD)算法。B2PD利用动作价值先验来引导条件变分自编码器(CVAE)生成一个高价值行为支持集。由此产生的专家行为先验进一步蒸馏到智能体中,有效减少低效探索并实现稳定的策略优化,同时建立双向知识流机制。在基于状态和基于像素的任务上的实证评估验证了B2PD在保持稳定策略优化的同时显著提高了样本效率。更广泛地说,这项工作表明,在线学习过程中强制执行高质量的行为支持能有效缓解评论家引起的误差放大,使结构化行为先验能够以原则性和样本高效的方式指导策略更新。
英文摘要
Online reinforcement learning (RL) algorithms frequently exhibit poor sample efficiency and unstable learning dynamics, stemming from systematic critic estimation errors that are exacerbated by greedy policy updates. Existing behavior-prior reinforcement learning methods attempt to alleviate this issue by relying on offline pre-training to learn behavior models from fixed datasets and using policy priors to constrain online policy updates. However, the limited quality of offline datasets often hinders the ability to provide high-value policies that can effectively guide policy updates. The absence of expert trajectories significantly impairs online policy learning, leading to low sample efficiency and suboptimal performance. To address these challenges, we depart from conventional behavior prior approaches and propose a Bidirectional Behavior Prior Distillation (B2PD) algorithm. B2PD leverages action-value priors to guide a conditional variational autoencoder (CVAE) in generating a high-value behavior support set. The resulting expert behavior priors are further distilled into the agent, effectively reducing inefficient exploration and enabling stable policy optimization, while establishing a bidirectional knowledge flow mechanism. Empirical evaluations on both state- and pixel-based tasks verify that B2PD substantially improves sample efficiency while maintaining stable policy optimization. More broadly, this work shows that enforcing high-quality behavioral support during online learning effectively mitigates critic-induced error amplification, enabling structured behavior priors to guide policy updates in a principled and sample-efficient manner.