arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08726cs.LGcs.AI

PAST:基于完整学生轨迹的特权适配用于在线策略自蒸馏

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

Yangyang Feng, Zhuoyan Feng, Junlan Chen

首次发表
浏览论文内容

中文总结 AI 辅助

PAST将完整学生轨迹作为OPSD教师的额外特权信息,通过前向KL蒸馏等方法优化教师,在三个数学推理基准上使Avg@12宏观平均值较Vanilla OPSD提升5.6个百分点。

中文摘要 AI 辅助

在线策略自蒸馏(OPSD)使用特权教师监督推理模型,该模型基于自身rollout采样的前缀进行学习。然而,每次rollout还会揭示学生响应的展开过程及是否成功,这是标准OPSD未用于构建教师的学生特定事后信息。我们引入基于学生轨迹的特权适配(PAST),将每个完整学生轨迹视为OPSD教师的额外特权信息,同时保留学生的蒸馏前缀不变。PAST在正确轨迹上保留学生的下一个token分布,并利用失败轨迹在学生邻近正则化下使教师向经验证的成功适配。我们表征了这种轨迹条件教师可向仅前缀学生传递的内容:前向KL蒸馏将教师分布投影到给定前缀的条件算术均值,该投影区分了仍为特权的轨迹特定变异与学生可获得的均值策略偏移。对于正确轨迹,未裁剪的总体目标还将冻结学生作为理想分布不动点。在三个数学推理基准上,PAST使Avg@12宏观平均值较Vanilla OPSD提升5.6个百分点。2×2因子研究显示,完整轨迹访问与教师适配均带来增益,而轨迹移除与打乱实验证实,适配后的教师使用匹配的事后上下文。

英文摘要

On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollout also reveals how the student's response unfolds and whether it succeeds, student-specific hindsight that standard OPSD does not use to form the teacher. We introduce Privileged Adaptation from Student Trajectories (PAST), which treats each completed student trajectory as additional privileged information for the OPSD teacher while leaving the student's distillation prefixes unchanged. PAST preserves the student's next-token distribution on correct trajectories and uses failed trajectories to adapt the teacher toward verified success under student-proximity regularization. We characterize what such a trajectory-conditioned teacher can transfer to a prefix-only student. Forward-KL distillation projects the teacher distributions to their conditional arithmetic mean given the prefix. This projection separates trajectory-specific variation that remains privileged from the mean policy shift available to the student. For correct trajectories, the unclipped population objective also has the frozen student as an ideal distributional fixed point. Across three mathematical reasoning benchmarks, PAST improves the Avg@12 macro average over Vanilla OPSD by 5.6 percentage points. A $2\times2$ factorial study shows gains from both complete-trajectory access and teacher adaptation, while trajectory removal and shuffling confirm that the adapted teacher uses the matching hindsight context.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑