arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弱到强泛化的在策略反向蒸馏

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun

arXiv 2609.08798首次发表:更新:

发表机构

KAIST AI; Microsoft; University of Toronto; Mila; Université de Montréal; CIFAR AI Chair(韩国科学技术院人工智能学院; 微软公司; 多伦多大学; 米拉计算科学研究所; 蒙特利尔大学; 加拿大高级研究院人工智能主席项目机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出在策略反向蒸馏(OPRD),通过评估教师策略偏移并放大验证器驱动的梯度分量,在弱到强泛化中加速学生优化并超越教师,且适用于多教师和强到弱蒸馏场景。

AI 中文摘要

弱到强泛化探讨的是,更强的模型能否从较弱的监督者那里学习并超越它们。这一问题对于连续模型代际更迭和多领域整合尤为重要,因为在这些场景中,从头开始重复进行前沿规模的训练后处理可能代价高昂。然而,传统的蒸馏方法将弱教师视为优化目标,可能将其能力上限强加于学生模型。我们提出了在策略反向蒸馏(OPRD),该方法在学生模型生成的回放轨迹上评估教师策略相对于其参考策略的策略偏移,并沿着该方向放大学生模型中由验证器驱动的策略梯度分量。通过仅重新缩放由验证器支持的更新,OPRD 保留了策略优化的平稳点,同时加速了超越教师的学习进程。在连续模型迁移和多教师蒸馏两种场景中,OPRD 相比现有的强化学习和蒸馏方法,以更少的学生更新次数取得了更高的性能。响应风格分析表明,OPRD 学生模型更接近仅使用基于验证器的强化学习训练的模型,而非其弱教师模型,这表明教师指导加速而非改变了学生自身的优化方向。在传统的强到弱蒸馏中的结果进一步表明,无论能力排序如何,OPRD 都能有效地将验证器驱动的策略优化与教师指导相结合。

英文摘要

Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.

Comments38 pages, 18 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑