arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06296cs.LG

无任何监督的在线策略自蒸馏

On-Policy Self-Distillation without Any Supervision

Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出无监督在线策略自蒸馏(U-OPSD),仅用模型自身生成结果实现在线策略自蒸馏,在多数学基准上优于基础模型,部分场景超越OPSD、GRPO等监督方法。

中文摘要 AI 辅助

在线策略(自)蒸馏(OPD / OPSD)在大型语言模型(LLM)的后训练阶段展现出强大潜力。然而,现有方法仍严重依赖外部监督,包括真值信号、环境反馈或更大模型的指导,因此未能实现真正的“自”蒸馏。本研究表明,仅利用模型自身生成结果并通过内部一致性即可实现在线策略自蒸馏,我们提出了无监督在线策略自蒸馏(U-OPSD)。U-OPSD首先采样多个rollout结果,在自一致性阈值下通过多数投票构建伪解;随后将教师分布基于最短伪解进行条件设置,并将其蒸馏到模型最长错误补全的前缀中,使模型能在自身确信错误的位置进行精准修正。在不同基准、基础模型和训练设置下,U-OPSD始终优于基础模型,且与带真值(GT)的监督方法(如OPSD和GRPO)相当或超越。在AIME24、AIME25、HMMT25、MATH500和AMC23基准上,Qwen3非思考模式下,U-OPSD在4B和8B规模时较基础模型分别提升8.5%和10.7%,较OPSD平均提升3.2%和2.3%;在思考模式下,U-OPSD与OPSD表现相当,4B规模时超越OPSD 0.9%,8B规模时与OPSD持平,同时分别超越GRPO 0.7%和1.1%。

英文摘要

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We propose unsupervised on-policy self-distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseudo solution by majority vote under a self-consistency threshold. It then conditions the model's distribution on the pseudo-solution and distills itself on the disagreeing completions, allowing the model to correct itself precisely where it is confidently wrong. Across diverse benchmarks, base models, and training settings, U-OPSD consistently improves over the base models and matches or surpasses supervised methods with ground truth (GT) such as OPSD and GRPO. On five mathematical reasoning benchmarks, i.e., AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD improves over the base model by 8.5% and 10.7% on Qwen3 non-thinking mode at 4B and 8B scales, and outperforms OPSD by 3.2% and 2.3% on average, respectively. In thinking mode, U-OPSD stays on par with OPSD, ahead by 0.9% at 4B and level at 8B and surpassing GRPO by 0.7% and 1.1%, respectively. Code is available at [https://github.com/williamium3000/u-opsd](https://github.com/williamium3000/u-opsd).

发表机构

  • UC San Diego(加州大学圣迭戈分校)
  • Georgia Institute of Technology(佐治亚理工学院)
  • University of Maryland, College Park(马里兰大学帕克分校)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑