arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10473cs.LGcs.AI

用于高效在线强化学习微调的无评判者预训练

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Daoyi Li, Yixian Zhang, Wenbo Ding, Yu Wang, Chao Yu

AI总结:

该研究针对离线转在线强化学习中复用离线评判者导致的适配问题,提出无评判者预训练范式,兼容主流算法且在多任务上表现更优。

AI中文摘要:

离线转在线(O2O)强化学习旨在利用在静态数据集上预训练的策略,同时通过在线交互对其进行改进。然而,直接复用离线训练的评判者(critic)会阻碍在线微调:随着策略和数据分布快速变化,从离线训练继承的价值估计可能与在线环境失配,导致策略改进不准确、探索效率低下。为解决该问题,我们提出无评判者预训练(Critic-Free Pretraining,CFP):一种完全摒弃离线评判者训练方法的高效范式,允许全新初始化的评判者适配,无需继承有偏估计。CFP 可兼容多种主流 O2O 算法,在各类任务上始终与传统 O2O 算法表现相当或更优,在多项挑战性任务上增益尤为显著。

英文摘要:

Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online fine-tuning: as the policy and data distribution change rapidly, value estimates inherited from offline training may become misaligned with the online environment, leading to inaccurate policy improvement and inefficient exploration. To address this problem, we introduce Critic-Free Pretraining: an efficient paradigm that completely abandons the approach of offline critic training, allowing a freshly initialized critic to adapt without inheriting biased estimates. CFP is compatible with various mainstream O2O algorithms and consistently matches or improves upon conventional O2O algorithms across a diverse set of tasks, with particularly pronounced gains on several challenging tasks.

↑