策略内参数更新方向是LLM后训练泛化能力的基础
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
浏览论文内容
中文总结 AI 辅助
本研究提出OPSFT方法,将策略内训练识别的参数更新方向约束至SFT,从而将策略内范式的泛化优势迁移至SFT,兼顾高训练效率与高质量轨迹利用。
中文摘要 AI 辅助
策略内(on-policy)后训练范式的强大泛化性能激发了对参数更新行为的研究。然而,这些研究仅将观察到的行为视为策略内训练的副产品,忽视了其作为优化原则以改善其他范式(如监督微调(SFT))泛化能力的潜力。为解决这一局限,我们探究是否存在一种特定的策略内更新行为能够实现上述改进。首先,我们的分析揭示,SFT沿一致方向更新参数,而策略内范式在训练过程中持续调整方向。这一差异促使我们关注每个参数的累积更新方向,将其视为一种有前景的行为。随后,我们通过提出策略内方向约束监督微调(OPSFT),评估其在改善泛化方面的有效性。OPSFT将SFT的更新约束至由策略内范式识别的方向上。OPSFT的强劲性能表明,策略内范式的泛化优势可通过参数更新方向迁移至SFT。一旦识别出该方向,即使SFT将其更新约束于此方向,也能实现泛化。这一发现通过将策略内范式的强泛化能力与SFT的优势(包括高训练效率和利用高质量轨迹的能力)相结合,带来两项实际益处。在效率方面,我们使用少量策略内训练步骤识别支持强泛化的更新方向,随后应用OPSFT以实现高训练效率。在利用高质量轨迹方面,OPSFT可利用这些轨迹沿其更新方向继续改进后训练模型,而不会破坏从策略内训练中学到的能力。
英文摘要
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.
发表机构
- State Key Lab. of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全国家重点实验室)
- University of Chinese Academy of Sciences(中国科学院大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。