发表机构
University of Southern California; Intel AI(南加州大学; 英特尔人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出在线策略功率蒸馏(OPPD),通过序列蒙特卡洛采样和教师功率分布加权,训练模型单次生成即可实现锐化采样效果,显著提升数学推理准确率,并与GRPO互补。
AI 中文摘要
语言模型分配给正确答案的概率可能高于任何单个错误答案,但模型仍然通常会采样出一个错误答案,因为所有错误答案合计持有更高的概率。功率分布将每个完整答案的概率提升至大于1的幂次并重新归一化,从而将概率向模型认为最可能的答案倾斜(即锐化)。从该分布中采样可在不改变参数的情况下提升推理能力,但每个查询需要大量带评分的候选答案。我们证明,模型可以通过训练在单次生成中产生此类答案。在线策略功率蒸馏(OPPD)运行一个序列蒙特卡洛采样器,其中被训练的模型生成候选答案,而冻结的教师模型的功率分布为其加权;相同的概率在最大似然更新中为每个答案加权。训练使单次生成准确率在MATH500上比未训练模型在同一温度下最高提升23.0个百分点,在GSM8K上提升27.3个百分点;单次生成得分比已发表的64个候选答案的功率采样高2.4和3.5个百分点,恢复了16个候选答案给未训练模型带来的增益的94%。作为对比,针对使用来自相同检查点和预算的验证奖励训练的GRPO,OPPD在MATH500、GSM8K和AIME上分别高出3.8、4.0和5.4个百分点,且不使用参考答案;两者互补,在GRPO之后应用OPPD可额外增加最多9.3个百分点。仅使用数学数据训练,OPPD将HumanEval准确率提升最多5.3个百分点。一个损失系数使模型吸收的锐化指数在1.19到2.02之间变化,而普通在线策略蒸馏为1.14,且该指数主要基于模型自身的答案上升。增益在不同模型家族和规模中保持,包括已经使用验证奖励训练的模型,其中降低温度无效果,而OPPD在MATH500上增加4.4个百分点。代码:此https URL。
英文摘要
A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but needs many scored candidates per query. We show that a model can instead be trained to produce such answers in one generation. On-policy power distillation (OPPD) runs a sequential Monte Carlo sampler in which the model being trained generates candidates and a frozen teacher's power distribution weights them; the same probabilities weight each answer in a maximum-likelihood update. Training raises single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature, and one generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94 percent of the gain that 16 candidates give the untrained model. For context, against GRPO trained with verified rewards from the same checkpoint and budget, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME using no reference answers; the two are complementary, and OPPD applied after GRPO adds up to 9.3 points. Trained only on mathematics, OPPD raises HumanEval accuracy by up to 5.3 points. One loss coefficient moves the sharpening exponent the model absorbs between 1.19 and 2.02, against 1.14 for ordinary on-policy distillation, and it rises mostly on the model's own answers. Gains hold across model families and sizes, including a model already trained with verified rewards, where lowering the temperature gives nothing and OPPD adds 4.4 points on MATH500. Code: https://github.com/ArminAzizi98/OPPD.