arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33127cs.LGcs.AI

策略可塑性在离线到在线强化学习中至关重要:为在线适应重新拟合离线策略

Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation

Yuheng Huang, Yunpeng Qing, Yixiao Chi, Yilun Kong, Changqing Zou

首次发表
浏览论文内容

中文总结 AI 辅助

针对离线到在线强化学习中离线策略可塑性下降的问题,提出REFIT方法,通过蒸馏到新初始化的学生网络并冻结随机子集来恢复可塑性,在D4RL和OGBench上超越现有方法。

中文摘要 AI 辅助

离线到在线强化学习(O2O RL)已成为一种实用范式,它利用静态离线数据集预训练策略,随后通过在线交互来调整策略。现有的O2O方法主要通过价值校准来处理这一过渡,通常将离线训练的策略视为给定的初始化。我们则从网络可塑性的角度研究O2O适应,探讨离线训练的策略是否仍能充分适应在线学习。受控实验表明,即使在离线性能基本饱和之后,在静态离线数据上的长时间优化也会逐渐降低网络可塑性,且较低的可塑性与较弱的后续在线改进相关。基于这些观察,我们提出了通过新鲜初始化和策略迁移恢复可塑性(REFIT),这是一种用于O2O过渡的轻量级模型级方法。在在线微调之前,REFIT将离线策略蒸馏到一个新初始化的学生网络中,同时暂时冻结学生单元的一个随机子集,将学到的离线行为迁移到一个更具可塑性的策略初始化上。在D4RL和OGBench上的大量实验表明,REFIT在Cal-QL和IQL两种骨干网络上均持续取得比现有O2O即插即用方法更高的总体性能,而可塑性诊断和消融实验进一步提供了网络可塑性恢复的证据。

英文摘要

Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while generally treating the offline-trained policy as a given initialization. We instead study O2O adaptation from the perspective of network plasticity, asking whether the offline-trained policy remains sufficiently adaptable for online learning. Controlled experiments show that prolonged optimization on static offline data progressively reduces network plasticity even after offline performance has largely saturated, and that lower plasticity is associated with weaker subsequent online improvement. Motivated by these observations, we propose REstoring plasticity via Fresh Initialization and policy Transfer (REFIT), a lightweight model-level method for the O2O transition. Before online fine-tuning, REFIT distills the offline policy into a freshly initialized student while temporarily freezing a random subset of student units, transferring the learned offline behavior to a more plastic policy initialization. Extensive experiments on D4RL and OGBench demonstrate that REFIT consistently achieves higher aggregate performance than existing O2O plug-in methods across both Cal-QL and IQL backbones, while plasticity diagnostics and ablations provide further evidence of restored network plasticity.

↑