发表机构
Shanghai University of Finance and Economics; Meituan; The Chinese University of Hong Kong, Shenzhen; Peking University(上海财经大学; 美团; 香港中文大学(深圳); 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SAF框架解决RLVR与OPD固定系数融合的熵崩溃问题,通过四阶段机制优化优势融合,在7个推理与代码生成基准上提升模型性能与训练稳定性。
AI 中文摘要
带可验证奖励的强化学习(RLVR)向每个token广播单一的响应级奖励,而策略内蒸馏(OPD)会针对更强的教师模型对每个token打分以获取密集优势,但性能上限为教师模型的质量,且不鼓励超出该质量的探索。二者的互补性使结合RLVR与OPD颇具前景,但我们发现用固定系数融合两种优势会因两种校准误差引发熵崩溃:一是幅度不匹配,token级OPD优势可能远超有界RLVR优势并抹去其信号;二是时间不匹配,持续全强度的OPD会不断将学生模型拉向教师模型,限制了超越教师所需的探索。我们提出SAF(稳定优势融合)框架,通过仅应用于OPD优势的轻量四阶段流水线解决上述问题:用于幅度控制的先稀疏后压缩机制,以及用于时间控制的先预热后退火机制,每个阶段可独立切换且仅增加可忽略的开销。以GRPO实例化RLVR,我们在7个数学推理与代码生成基准上用Qwen3-1.7B/4B/8B评估SAF:SAF避免了熵崩溃,且始终优于固定系数的GRPO+OPD融合,在全部6个模型-领域设置中综合得分提升0.51-2.70%,同时实现更稳定的训练。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.
CommentsWorking in progress