arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应专家引导用于高效在线策略强化学习

Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning

Daniele Affinita, Ming Xu, Rudolf Reiter, Davide Scaramuzza, Pascal Fua

arXiv 2610.06019首次发表:更新:

发表机构

EPFL; University of Zurich(洛桑联邦理工学院; 苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种在线策略强化学习方法,通过可学习参数动态调整专家与学习者的控制份额,在训练早期利用次优专家提升样本效率,随能力增强自动退出,最终优于专家。

AI 中文摘要

随着大规模并行模拟的出现,诸如PPO之类的在线策略强化学习方法已成为许多领域中的标准方法。然而,从零开始学习在样本效率上较低,并且未能利用可能存在的次优专家,例如启发式方法、基于模型的控制器或在相关任务上训练的策略。这样的专家通常可用,并能指导早期训练,但其次优性限制了最终性能。挑战因此变为在专家引导与从奖励中学习之间取得平衡。现有方法通过混合权重、调度或评估驱动的课程来设置专家的影响力。或者,它们通过额外的学习组件(如对专家动作的批评者或辅助代理)来调整专家影响力。然而,没有一种方法使用与策略本身相同的在线策略目标来优化专家影响力。我们提出了一种方法,其中学习者和专家在每个训练回合内交替控制,专家的控制份额是一个单一的可学习参数,与策略联合优化。学习者在训练早期受益于专家,但随着学习者能力的增强,专家的控制份额下降,最终完全消失。这使得学习者独自行动,并且优于次优专家。我们在两个基准的34个任务上评估了我们的方法,涵盖离散和连续动作空间,使用学习型和基于模型的专家。我们的方法在引导和未引导基线上提高了样本效率,同时需要最小的超参数变化。专家的份额随着学习者的改进而衰减至零,在专家不再有用时消失。

英文摘要

With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.

Comments21 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑