发表机构
Shanxi University; Nanjing University; The Chinese University of Hong Kong; Tianjin University(山西大学; 南京大学; 香港中文大学; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次将未训练的决策模型Jev作为参考策略、探索判断器和回放评分器用于RL训练,在9个MiniGrid任务和3款Atari游戏上提升了RL的样本效率与性能。
AI 中文摘要
基础模型为强化学习(RL)提供先验,缓解其长期存在的样本效率和迁移性弱点,但它们的逐token生成方式使查询具有序列性且成本高昂。Jev是最近发布的一种决策模型,无需生成内容,仅通过一次前向传播即可返回经过校准的类型化答案。现有研究将强化学习中的基础模型要么作为待训练模型,要么作为待提示的生成器,而Jev不属于这两类,迄今为止仅在单一领域中作为黑盒使用。因此,这类模型在RL环境中自行决策的效果,以及它作为训练组件如何改进RL,仍未得到解决。为此,本文首先研究RL系统的对象对其使用的答案提出的要求,确定Jev可满足所有要求,仅不满足价值函数的核心用途。其余对象形成的位置各可承担多种角色。随后,我们构建算法,在其中三个位置使用Jev:作为参考策略、探索判断器和回放评分器,以提升样本效率、探索能力和学习性能。在9个MiniGrid任务和3款Atari游戏上,使用Jev进行训练的表现优于标准RL学习者,包括在学习者单独无法取得进展的场景中,且模型本身保持未训练状态。据我们所知,本文首次将Jev用于RL学习过程,并证实冻结的决策模型可作为RL训练的可用组件,为进一步探索Jev及其他先进决策模型如何改进RL提供了方向。
英文摘要
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements that the objects of an RL system place on the answers they consume, and establish that Jev can fulfill all of them except the cardinal use of a value function. The remaining objects form positions that admit several roles each. We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance. Across nine MiniGrid tasks and three Atari games, training with Jev outperforms a standard RL learner, including where the learner makes no progress alone, while the model itself remains untrained. To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.