arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VISTA:面向在线自蒸馏的验证器驱动的学生到教师适配

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua, Haoyu Cao, Linli Xu

arXiv 2608.28306首次发表:更新:

发表机构

University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence(中国科学技术大学; 认知智能国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VISTA 复用标准 OPSD 的 rollout 与损失函数,通过结果验证的 rollout 将教师适配学生分布,在多基准数据集上提升了不同规模 Qwen3 模型的 Avg@12。

AI 中文摘要

在线自蒸馏(OPSD)通过使用仅问题的学生模型在其自身的 rollout 上进行训练来提升推理能力,其中学生模型使用来自特权教师的密集 token 级监督,而该特权教师还能看到参考解决方案。然而,标准 OPSD 将教师分布视为学生 rollout 过程中的固定目标,仅更新学生模型,尽管特权条件并不能保证教师总能为仅问题推理提供最合适的目标。因此,当教师分布与有效的学生推理不一致时,这种单向监督可能会误导学生。为此,我们提出了验证器驱动的学生到教师适配(VISTA),它保留了标准 OPSD 的学生更新,同时使用结果验证的 rollout 来使教师适配学生分布。在每次验证的 rollout 中,VISTA 还将此适配限制在教师与学生 KL 散度最大的前 k 个位置。值得注意的是,VISTA 复用了标准 OPSD 的 rollout 和损失函数,未引入额外采样或单独的奖励目标。在使用 Qwen3 模型(17 亿、40 亿和 80 亿参数)的 AIME24、AIME25 和 HMMT25 数据集上,VISTA 在所有规模下均取得最高的 Avg@12,分别比 OPSD 提升了 0.6、0.7 和 2.1 个点。这些结果证明了来自结果验证 rollout 的学生监督的价值,并强调学生到教师适配是 OPSD 的一个有前景的方向。

英文摘要

On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, standard OPSD treats the teacher distribution as a fixed target along the student's rollout and updates only the student %, although -- even though privileged conditioning does not guarantee that the teacher always provides the most appropriate target for problem-only reasoning. This one-way supervision can therefore misdirect the student when the teacher distribution is misaligned with valid student reasoning. We therefore introduce Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution. Within each verified rollout, VISTA further restricts this adaptation to the top-$k$ positions with the largest teacher--student KL divergence. Notably, VISTA reuses the rollout and loss function from standard OPSD, introducing no additional sampling or separate reward objective. Across AIME24, AIME25, and HMMT25 with Qwen3 models at 1.7B, 4B, and 8B, VISTA achieves the highest Avg@12 at every scale, improving over OPSD by $0.6$, $0.7$, and $2.1$ points, respectively. These results demonstrate the value of student supervision from outcome-verified rollouts and highlight student-to-teacher adaptation as a promising direction for OPSD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑