发表机构
Tsinghua University; Alibaba Group(清华大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究自生成与奖励加权数据上的策略微调,提出统一理论分析RE(S)算法,证明其全局收敛与Θ(1/T)速率,并揭示适当离策略更新可避免局部陷阱、加速收敛。
AI 中文摘要
我们研究了在自生成和奖励加权数据上微调策略模型的学习动态,特别关注REINFORCE的一种广义版本——记为RE(S)——该版本每S≥1个梯度步骤更新一次 rollout 分布。先前在 bandits 和强化学习中的工作为策略梯度方法建立了丰富的理论,并且 on-policy 采样(即较小的S,理想情况下为1)常被视为其成功的关键;然而,在诸如大语言模型的后训练等突出应用中,即使 rollout 分布更新不频繁,奖励引导的自训练也被证明是有效的,但这些 off-policy 方法的收敛性质的理论理解仍然有限。为弥合这些差距,我们为RE(S)开发了一个统一理论,覆盖S≥1的全部范围:它可被解释为一个分阶段优化过程,其中每个阶段采取S个梯度步骤来最小化到固定奖励加权 rollout 分布的 Kullback-Leibler 距离。对于具有 softmax 策略的多臂老虎机,我们的深入分析和数值实验揭示了三个关键发现:(1)对于任何固定的S,随着 rollout 分布更新次数 B = ⌊T / S⌋ → ∞(其中T表示梯度步骤数),RE(S) 全局收敛到最优策略;(2)我们证明了紧的上下界,表明RE(S)的次优性差距达到渐近 Θ(1 / T) 收敛速率,而S仅影响 burn-in 阶段的长度;(3)当以较小的最优动作概率初始化于弱策略时,RE(1) 会在次优策略附近长时间陷入,而具有合适S的RE(S)避免了绕路并显著更快地收敛到全局最优,突出了离策略性在此情况下的优势。
英文摘要
We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small $S$, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of $S \ge 1$: it can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed $S$, RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates $B = \lfloor T / S \rfloor \rightarrow \infty$, where $T$ denotes the number of gradient steps; (2) we prove tight two-sided bounds showing that the suboptimality gap of RE(S) achieves an asymptotic $Θ(1 / T)$ convergence rate, while $S$ only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable $S$ avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.