MaD-RL:通过强化学习匹配分布以校准大语言模型
MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对强化学习后训练导致输出多样性降低的问题,提出分布匹配框架MaD-RL,通过匹配潜在属性分布到目标分布,并设计KL和JS散度奖励函数,在数学推理和编程任务中验证有效性。
中文摘要 AI 辅助
强化学习(RL)广泛用于语言模型的后训练,以最大化分配给单个模型输出的奖励,例如来自二元验证器的分数或基于人类反馈训练的奖励模型。然而,诸如合成数据生成、公平性相关的约束满足以及策略探索等应用,需要控制模型生成输出的分布,而不仅仅是最大化期望奖励。我们提出了一个通用的基于RL的框架,用于\textit{分布匹配},允许将模型输出的潜在分类属性的分布匹配到指定的目标分布。实证上,我们证明了诸如组相对策略优化(GRPO)等主流后训练方法通过将策略概率集中到单一模式,降低了输出多样性。熵正则化和采样温度可以改善分布的扩散,但效果有限,仅适用于词元空间且朝向均匀分布。我们表明,该领域的先前工作是涉及$L_2$散度的分布匹配的特例。然后,我们为其他散度(如KL和Jensen-Shannon)提出了奖励函数,并给出了理论上的合理性论证。最后,我们在涉及数学推理和编程的一组实验中展示了我们方法的有效性。
英文摘要
Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data generation, fairness-related constraint satisfaction, and policy exploration require controlling the distribution of outputs across model generations rather than only maximizing expected reward. We propose a general RL-based framework for \textit{Distribution Matching} allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution. Empirically, we demonstrate that dominant post-training recipes such as Group Relative Policy Optimization (GRPO) reduce output diversity by concentrating policy probability towards a single mode. Entropy regularization and sampling temperature can improve the spread of the distribution but have constrained effectiveness, limited to apply only in token space and toward uniform distributions. We show that prior work in this area is a specific case of Distribution Matching involving the $L_2$ divergence. We then propose reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification. Finally, we demonstrate the effectiveness of our approach on a set of experiments involving mathematical reasoning and programming.