克服大语言模型推理中同策略自蒸馏的扩展限制
Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
浏览论文内容
中文总结 AI 辅助
针对同策略自蒸馏在模型扩展时效果下降的问题,提出OASIS方法,通过监督经验证的同策略轨迹并利用未经验证的模型生成作为教师上下文,在多个基准上显著提升推理性能。
中文摘要 AI 辅助
同策略自蒸馏(OPSD)训练学生模型沿其自身采样的轨迹匹配特权教师分布。标准OPSD将这种监督应用于未经验证的学生轨迹,同时以特权上下文(通常是参考解决方案)为条件来指导教师。我们在因子分析中分离这些角色,发现脚手架正确性对下游准确率的影响大于上下文正确性。未经验证的脚手架会产生模仿差距,因为教师可以使用学生无法获得的信息。这种差距随模型规模增大而缩小,但OPSD仍主要监督未经验证的轨迹。相比之下,即使教师以学生自身不成功的轨迹为条件,经验证的脚手架仍保持有效。基于这一发现,我们引入OASIS,它保留OPSD目标,但主要监督经标签验证的同策略轨迹,并用未经验证的模型生成尝试取代书面解决方案作为教师上下文。因此,OASIS仅需最终答案标签。在AIME 2024、AIME 2025和HMMT 2025上,针对Qwen3-1.7B、4B和8B,OASIS平均比基础模型提升3.2-3.8分,而OPSD的增益从1.7B时的3.05分降至8B时的0.14分。在8B规模下,OASIS比OPSD提升3.05分,表明经验证的同策略脚手架在模型扩展时保持了自蒸馏的有效性。
英文摘要
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.
发表机构
- North South University(南北大学)
- KAIST(韩国科学技术院)
- Andria Labs
机构由 AI 辅助整理,请以论文原文为准。