arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线RL微调真的需要预训练Q函数吗?

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Finn

arXiv 2607.27203首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对在线RL微调是否需预训练Q函数的问题,发现预训练Q函数增益有限,提出IPE方法,在连续控制基准中使微调性能平均提升1.26倍。

AI 中文摘要

预训练后微调已成为学习高性能策略的主流方法,在基于值的强化学习(RL)中,这引发了一个自然问题:给定预训练策略,是否也应在离线数据上预训练Q函数?传统观点认为应当如此,但近期结果显示,采用随机初始化Q函数的在线RL无需预训练Q函数即可得到高性能且可靠的策略。本文系统研究在预训练基础策略上微调时,预训练Q函数是否真有帮助,结果意外发现,简单的Q函数预训练通常相比随机初始化几乎无增益。我们表明这源于根本不匹配:预训练期间学习的Q函数针对预训练策略的Q函数,而非在线微调收敛的Q函数,该差距即使经离线值最大化仍存在。基于此发现,我们提出策略集成初始化(IPE),这是一种训练多个多样化策略并利用其汇集的rollout来引导在线RL中Q函数学习的简单方法。在一系列具有挑战性的连续控制基准测试中,IPE相比简单的Q函数预训练,微调性能平均提升1.26倍。

英文摘要

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function. In this paper, we systematically study whether pretraining the Q-function actually helps when fine-tuning on top of a pretrained base policy. We find, surprisingly, that naive Q-function pretraining often provides little benefit over random initialization. We show this stems from a fundamental mismatch: the Q-function learned during pretraining targets the pretrained policy's Q-function, not the Q-function that online fine-tuning converges to, and this gap persists even after offline value maximization. Motivated by this finding, we propose Initialization via Policy Ensemble (IPE), a simple method that trains multiple diverse policies and uses their pooled rollouts to bootstrap the Q-function learning in online RL. Across a suite of challenging continuous control benchmarks, IPE yields an average 1.26x improvement in fine-tuning performance over naive Q-function pre-training.

CommentsFixed typo in author name

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑