arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09745cs.LGcs.AIstat.ML

SR-OPSD:自引用同策略自蒸馏

SR-OPSD: Self-Referenced On-Policy Self-Distillation

  • Shanghai University of Finance and Economics(上海财经大学)
  • Imperial College London(伦敦帝国学院)
  • University of Science and Technology of China(中国科学技术大学)
  • University College London(伦敦大学学院)
  • Nanyang Technological University(南洋理工大学)
  • Technical University of Denmark(丹麦技术大学)
  • University of Copenhagen(哥本哈根大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan, Kaiyu Li, Che Liu, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Li Zeng

AI总结:

该研究针对同策略自蒸馏的优化不稳定问题,提出SR-OPSD方法,通过几何插值结合Rényi散度优化蒸馏目标,在多任务大语言模型实验中取得最优或竞争性能。

AI中文摘要:

同策略自蒸馏(OPSD)将反馈转化为待优化策略生成轨迹上的稠密 token 级监督,可有效补充具有稀疏结果奖励的强化学习。但 OPSD 中使用的自教师策略通常是策略的停止梯度或指数移动平均副本,且受额外上下文信息约束,会随学生策略及其同策略上下文分布共同演化,用固定投影目标匹配该移动目标易导致优化不稳定或分布过度集中。基于此,本文提出自引用同策略自蒸馏(SR-OPSD)。在固定学生生成的上下文时,token 级变分特征将有效蒸馏目标确定为自教师策略与参考策略的几何插值;同时,采用 Rényi 散度族对投影几何进行泛化。该公式将自适应目标的放置位置与学生向其投影的方式分离:插值系数控制底层目标,Rényi 阶控制投影几何及其对 token 级密度比的敏感性。在科学评估、数学推理和代码生成任务中,使用多个大语言模型开展的大量实验表明,SR-OPSD 在各类设置下均达到了当前最优或具有竞争力的性能。

英文摘要:

On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target--student probability mismatches translate into updates. We propose \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}, which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward Rényi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the Rényi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.

↑