arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准部分重置:防止连续强化学习中的策略崩溃

Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning

Luc McCutcheon, Evangelos Chatzaroulas, Saber Fallah

arXiv 2607.24996首次发表:更新:

AI 中文总结

研究连续强化学习中神经网络训练问题,提出校准部分重置(CPR)优化器,定期拉低效用神经元至初始化,拉动强度依效用缩放。该方法避免策略崩溃,在多基准测试中表现优,揭示可塑性与峰值性能权衡,为连续学习提供新方向。

AI 中文摘要

神经网络在训练过程中会受到累积休眠神经元和表达能力丧失的阻碍,尤其是在非平稳数据设置中,如连续监督学习和强化学习。近期,神经元重置被用于维持梯度流和恢复可塑性。然而,完全单元重新初始化往往会牺牲峰值性能并使训练不稳定,导致策略崩溃。为在不破坏训练稳定性的情况下保持可塑性,我们提出校准部分重置(CPR),这是一种优化器,它会定期将低效用神经元拉向其初始化状态,拉动强度由每个神经元的效用缩放。与二元重置方法不同,部分重置避免了脆性;与均匀衰减不同,校准效用缩放将调整集中在最需要的单元上。在比较的方法中,只有CPR在SlipperyAnt超过4亿步训练中避免了策略崩溃,并且在连续MetaWorld和连续MinAtar基准测试中优于基于衰减和重置的先前方法。消融实验揭示了可塑性和峰值性能之间的可调权衡,突出了效用缩放重新初始化作为连续学习的一个有前景的方向。

英文摘要

Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary data settings, such as continual supervised and reinforcement learning. Recently, neuron resets have been used to maintain gradient flow and restore plasticity. However, full unit reinitialization often sacrifices peak performance and can destabilize training, leading to policy collapse. To preserve plasticity without destabilizing training, we propose Calibrated Partial Resets (CPR), an optimizer that periodically pulls low-utility neurons toward their initialization, with pull strength scaled by each neuron's utility. Unlike binary reset methods, partial resets avoid brittleness; unlike uniform decay, calibrated utility-scaling concentrates adjustment on the units that need it most. Among compared methods, only CPR avoids policy collapse over 400M training steps in SlipperyAnt, and it outperforms prior decay and reset-based methods on Continual MetaWorld and Continual MinAtar benchmarks. Ablations reveal a tunable trade-off between plasticity and peak performance, highlighting utility-scaled reinitialization as a promising direction for continual learning.

CommentsRLC Continual RL Workshop (Oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑