arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07935cs.LG

面向策略内自蒸馏的自适应监督锚定

Adaptive Supervised Anchoring for On-Policy Self-Distillation

Meilin Yang, Zixuan Ding, Jianhao Nie, Weite Zhang, Yuxin Zhang, Zhiming Shao, Li Yu, Zhe Fu

首次发表
浏览论文内容

中文总结 AI 辅助

针对策略内自蒸馏的 rollout 条件信号退化问题,提出分离监督路径并自适应调整锚定强度的框架,提升了任务获取能力并优化了可塑性-稳定性权衡。

中文摘要 AI 辅助

策略内自蒸馏(OPSD)通过从学生模型采样的轨迹中蒸馏冻结教师模型的指导来调整语言模型,但其有效性关键依赖于这些轨迹的质量。研究表明,当学生的 rollout 偏离目标轨迹时,让教师模型基于偏离目标的前缀进行条件处理会大幅削弱其与任务相关的监督。受控前缀损坏实验揭示了这种失败模式,将其命名为 rollout 条件信号退化。为解决该问题,提出一种统一训练框架,分离两条互补监督路径:第一条保留 rollout 条件分布匹配,为学生实际访问的状态提供指导;第二条在规范真实上下文上应用监督交叉熵,避免将目标 token 强加于错误 rollout 前缀的不兼容性。采用 token 级 rollout-目标对齐来调整规范上下文锚的强度,在冷启动阶段强化该锚,随 rollout 质量提升而放松。在多个模型规模、两类任务族及通用推理基准上的实验显示,所提方法相较 OPSD 提升了任务获取能力,同时保留了通用能力,实现了更优的经验可塑性-稳定性权衡。这些发现确定了上下文质量是策略内自蒸馏的核心瓶颈,并证明了分离 rollout 条件指导与规范监督的价值。

英文摘要

On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depends critically on the quality of those trajectories. We show that when student rollouts drift from target trajectories, conditioning the teacher on off-target prefixes substantially weakens its task-relevant supervision. Controlled prefix-corruption experiments expose this failure mode, which we term rollout-conditioned signal degradation. To address this problem, we propose a unified training framework that separates two complementary supervision pathways. The first retains rollout-conditioned distribution matching, providing guidance on states the student actually visits. The second applies supervised cross-entropy on canonical ground-truth contexts, avoiding the incompatibility of imposing target tokens on erroneous rollout prefixes. Token-level rollout-target alignment is used to adapt the strength of the canonical-context anchor, emphasizing it during cold start and relaxing it as rollout quality improves. Experiments across multiple model scales, two task families, and general-reasoning benchmarks show that the proposed approach improves task acquisition over OPSD while preserving general capabilities, resulting in a more favorable empirical plasticity-stability trade-off. These findings identify context quality as a central bottleneck in on-policy self-distillation and demonstrate the value of separating rollout-conditioned guidance from canonical supervision.

发表机构

  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑