I-SDPO:实例级自适应自蒸馏策略优化
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
浏览论文内容
中文总结 AI 辅助
针对GRPO在全错推理组无有效相对信号的问题,提出I-SDPO,按实例路由选择特权自蒸馏或GRPO,在SciKnowEval上实现全领域最优,平均mean@16准确率提升至70.31%。
中文摘要 AI 辅助
组相对策略优化(GRPO)从一次推理组内的奖励差异中学习,但当所有采样响应均不正确时,无法获得有用的相对信号。特权自蒸馏可通过密集令牌监督填补这一空白,但在整个训练过程中应用它会产生另一种失败模式:教师模型是奖励目标的有偏低方差替代物,因此当策略能够生成成功轨迹后,持续模仿会阻碍奖励改进更新。我们引入I-SDPO(Instance-Level Adaptive Self-Distillation Policy Optimization,实例级自适应自蒸馏策略优化),将教师依赖视为与能力相关。I-SDPO对每个输入实例做出一次路由决策,并在该实例的推理组间共享:全错组使用特权自蒸馏目标,而任何成功组则保持GRPO不变。该设计仅在组相对奖励无信息时使用模仿。局部分析刻画了教师与奖励方向对齐的情况,并表明非零偏置蒸馏权重会诱导优化偏置下限。路由规则会随成功概率上升自动降低预期蒸馏率,无需人工设计的调度即可撤回教师影响。在SciKnowEval数据集上,I-SDPO在所有四个科学领域均取得最佳结果,将平均mean@16准确率从GRPO的56.67%提升至70.31%,最大领域增益达18.24个百分点。
英文摘要
Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with dense token supervision, yet applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate for the reward objective, so persistent imitation can oppose reward-improving updates after the policy becomes capable of producing successful trajectories. We introduce I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent. I-SDPO makes one routing decision per input instance and shares it across that instance's rollout group: all-incorrect groups use a privileged self-distillation objective, whereas any-success groups remain intact for GRPO. This design uses imitation only where group-relative rewards are uninformative. A local analysis characterizes when teacher and reward directions align and shows that a non-vanishing biased distillation weight induces an optimization bias floor. The routing rule automatically reduces the expected distillation rate as success probability rises, withdrawing teacher influence without a hand-designed schedule. On SciKnowEval, I-SDPO obtains the best result in all four scientific domains and improves average mean@16 accuracy from 56.67% with GRPO to 70.31%, with a maximum domain gain of 18.24 points.
发表机构
- Qwen Large Model Application Team, Alibaba(阿里通义千问大模型应用团队)
机构由 AI 辅助整理,请以论文原文为准。