arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04419cs.LGcs.AI

SPOT:面向在线策略蒸馏的稀疏探测与结果校准

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Zhengyu Chen, Yikun Ban, Shuang Qiu, Zhongxiang Dai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对在线策略蒸馏的缺陷,提出SPOT方法,通过获取-探索-利用程序优化探测与蒸馏,经多模型多基准实验验证,可提升推理性能并平衡解的质量与覆盖范围。

中文摘要 AI 辅助

在线策略蒸馏(OPD)为学生模型生成的轨迹提供了密集的教师监督,但标准反向KL训练可能会为其他合理的后续分配不足的概率。仅教师熵无法揭示不确定性是集中在少数合理的下一个标记上,还是分散在长长的概率尾部,也无法揭示学生是否已经很好地表示了这些候选。此外,局部教师概率可能无法预测下游成功。我们引入了面向在线策略蒸馏的稀疏探测与结果校准目标(SPOT),它解决了两个耦合决策:探测位置和蒸馏内容,通过获取-探索-利用程序实现。在获取阶段,位置级分数结合了归一化教师熵、小的top-k候选集捕获的概率质量以及学生-教师不匹配,以分配有限的探测预算。在探索阶段,SPOT通过验证器评分的学生后续评估教师提出的候选。在利用阶段,这些结果产生了一个闭式的KL正则化目标,该目标有利于具有更好下游结果的候选,同时保持锚定在教师分布上。在多个学生模型和推理基准上进行的大量实验证明了SPOT在提高推理性能的同时平衡解决方案质量和覆盖范围的有效性。

英文摘要

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-$k$ candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • East China Normal University(华东师范大学)
  • Beihang University(北京航空航天大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑