发表机构
Adelaide University; University of North Texas; Curtin University; The University of Sydney(阿德莱德大学; 北德克萨斯大学; 科廷大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自蒸馏中信用方向与幅度耦合问题,提出解耦信用自蒸馏(DCSD),通过信念边际探测和边际信息增益分别确定方向与幅度,在11个基准上取得最优,显著提升数学与多模态推理性能。
AI 中文摘要
RLVR提供了可靠的轨迹级信用,而OPSD则为令牌级信用提供了密集监督。这暴露了在利用教师监督更新步骤级信用方向和幅度时的根本耦合问题,阻碍了步骤获得可靠的信用方向和贡献幅度,同时使两者都易受教师判断错误和偏好差异的影响,我们的理论分析支持了这一观点。为了将信用方向与其贡献幅度分离,我们引入了解耦信用自蒸馏(DCSD),它在理论上将信用方向和幅度解耦为两个可靠信号,并用它们来校准特权教师监督。具体来说,我们设计了信念边际探测来确定信用方向,以及边际信息增益来量化信用幅度,从而实现从步骤到令牌的信用分配以进行策略优化。在11个基准测试中,DCSD在总体得分上优于GRPO、OPSD、RLSD和RLCSD。与基础模型相比,DCSD在数学推理上总体得分提高了8.45分,在多模态推理上提高了7.01分,同时纠正了6%令牌的信用方向,并将令牌信用幅度减少了1.5倍。
英文摘要
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.