发表机构
University of Washington; Allen Institute for AI(华盛顿大学; 艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究拆解语言模型的RL后训练算法,探究各因素对其结果的影响,为NLP领域人员提供RL后训练的入门指南。
AI 中文摘要
强化学习(RL)后训练已成为增强大型语言模型(LLM)能力的强大框架,使其具备出色的推理、数学和编码能力。然而对许多研究人员和从业者而言,经典RL背后的原理仍是一个“黑箱”。本研究对RL后训练算法进行拆解,探究每一步以阐明其实际运行机制。我们在受控且简化的环境中分离出带可验证奖励的RL机制,研究RL结果如何受基础模型的先验分布、奖励信号的粒度、提示分布的多样性以及模型规模的影响。我们以策略输出分布的熵为视角,对比预训练、SFT(监督微调)和RL后训练阶段学习到的分布,揭示每个阶段如何塑造模型的确定性。我们的研究阐明了这些选择如何相互作用影响后训练的成功,例如,所谓“虚假奖励”的效果取决于后训练所用的提示分布。我们还阐释了RL后训练的成功为何取决于基础模型是否已对期望行为分配足够的概率质量,这与RL中经典的探索概念相关联。最终,我们提供这份入门指南,供NLP领域希望将RL作为工具纳入其工具箱的人员使用。
英文摘要
Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers and practitioners, the principles behind classical RL remain a "black box". In this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface. By isolating the mechanics of RL with Verifiable Rewards in a controlled and simplified environment, we examine how RL outcomes are shaped by the base model's prior distribution, the granularity of the reward signal, the diversity of the prompt distribution, and model scale. We use the entropy of the policy's output distribution as a lens to compare the distributions learned through pretraining, SFT, and RL post-training, revealing how each stage shapes model certainty. Our investigation sheds light on how these choices interact to affect post-training success. For example, we show that the effect of so-called 'spurious rewards' depends on the prompt distribution used for post-training. We also provide insight into why the success of RL post-training depends on whether the base model already places sufficient probability mass on the desired behavior, linking it to the classical concept of exploration in RL. Ultimately, we provide this primer as a resource to those in the NLP community wishing to incorporate RL as a tool in their toolbox.
CommentsLink to website: https://minjang10.github.io/demystifying-rl-finetuning-web Link to code: https://github.com/sankarh-1/demystifying-rl-finetuning