arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励膨胀:强化学习的健康刺激

Reward Inflation: A Healthy Stimulus for Reinforcement Learning

Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang

arXiv 2610.02545首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出奖励膨胀,即在训练中逐步缩放奖励,作为强化学习的健康刺激,通过隐式近期加权和维持梯度信号提升适应性与可塑性,并在ALE和MuJoCo任务上验证其有效性,同时提出自适应变体Fed。

AI 中文摘要

在强化学习(RL)中,奖励作为主要的学习信号。然而,虽然奖励幅度在训练过程中通常保持固定,但其时间调制仍未得到充分探索。在本文中,我们提出奖励膨胀,即在训练过程中对奖励进行逐步缩放,并表明它可以作为RL的健康刺激。理论上,奖励膨胀引入了隐式的近期加权,在策略更新时上调近期转移的权重,从而实现更快的适应。我们进一步表明,通过维持梯度信号,当策略饱和时,奖励膨胀抑制了休眠神经元的出现,并有助于保持可塑性。在ALE游戏和MuJoCo任务上的实证结果证实了这些发现,表明适当水平的奖励膨胀有益于广泛的任务。最后,我们引入了Fed,一种自适应变体,可动态调整膨胀水平,并发现它通常优于固定膨胀。

英文摘要

Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑