arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02998cs.LGcs.AIcs.CL

验证后再蒸馏:用于在线蒸馏的提示级教师门控机制

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

  • AllSpark Team(AllSpark团队)

机构由 AI 辅助整理,请以论文原文为准。

Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan

AI总结:

该研究提出 TGOPD 方法,在提示级验证教师可靠性后再进行在线蒸馏,在数学、代码等任务的 4B 和 35B 规模模型上性能优于常规 OPD,还提升了教师 GPU 利用率。

AI中文摘要:

在线蒸馏(OPD)通过在学生模型的自身 rollout 中提供来自冻结教师模型的密集 token 级监督,加速了后训练过程。常规 OPD 会在所有提示上均匀应用这种监督,而不检查教师对每个提示是否可靠。由于反向 KL 具有模式寻求特性,自信但错误的教师可能会诱导出强但具有误导性的更新。分布代理(如熵或师生似然一致性)可衡量不确定性或一致性,但无法直接验证结果的正确性。我们引入教师门控在线蒸馏(TGOPD),其核心原则是:在允许密集监督前,需在提示级验证教师的可靠性。TGOPD 从少量由验证器评分的教师探测样本中估计可靠性,当可靠性检查通过时,将每个提示仅路由至密集 OPD;否则,路由至基于验证器的 GRPO。在数学、代码和指令跟随领域的 4B 和 35B 规模学生模型上,TGOPD 在全部 6 个单领域设置中均优于常规 OPD,且在多领域训练下,两种规模的七基准平均性能均更高。通过利用原本闲置的教师算力进行可靠性估计,TGOPD 还减少了异步 OPD 中教师侧的计算浪费,在实测的 4B 单领域运行中,教师节点的 GPU 利用率从 9.8%提升至 78.9%。

英文摘要:

On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.

补充信息

↑