arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15404cs.AI

谁教哪个词元?验证器门控多专家在线策略蒸馏用于科学推理

Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning

Xun Xu, Zaixi Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多教师在线策略蒸馏中教师信号稀疏异质的问题,提出验证器门控多专家在线策略蒸馏(VG-OPD),通过反事实增益授权、分歧定位和标准重要性加权,在科学推理基准上取得最佳性能。

中文摘要 AI 辅助

多教师在线策略蒸馏(OPD)正成为将专家能力整合到单一模型中的标准方法:先用强化学习训练专家,然后在其自身轨迹上将其蒸馏到学生模型中。现有方法在序列级别分配监督——每个提示分配给一个领域教师,每个词元获得相同的权重——这隐含地假设教师在整个回答中的有用性是均匀的。相反,我们发现有用的教师信号沿推理轨迹是稀疏且异质的,这引发了一个更精细的问题:谁应该教哪个词元?验证器门控多专家在线策略蒸馏(VG-OPD)通过验证来回答这一问题:专家在特定答案标准上的反事实增益授权该专家进行教学,其与学生模型的分歧定位监督,标准重要性设置其权重;门控KL作为加性词元级优势进入GRPO。在科学推理任务中,使用强化学习训练的能力专家实例化后,VG-OPD在4B和8B学生模型的七个基准上取得了最佳整体性能,在两个规模上均排名第一的有五个基准,在知识密集型科学推理任务上提升最大。进一步分析表明,性能提升来自定位已验证的监督,而非增加教师或蒸馏损失:错置相同的监督预算是最具破坏性的变化,而不加区分的蒸馏将强化学习拖至其自身下限之下,而门控蒸馏则将其提升。

英文摘要

Multi-teacher on-policy distillation (OPD) is becoming the standard way to integrate specialist capabilities into one model: train experts with RL, then distill them into the student on its own rollouts. Existing recipes assign supervision at the sequence level - each prompt goes to one domain teacher and every token receives the same weight - which implicitly assumes that a teacher is uniformly useful across a response. We find instead that useful teacher signal is sparse and heterogeneous along a reasoning trajectory, which raises a finer question: who should teach which token? Verifier-Gated Multi-Expert On-Policy Distillation (VG-OPD) answers it by verification: the counterfactual gain of an expert on a specific answer criterion licenses that expert to teach, its disagreement with the student localizes the supervision, and criterion importance sets its weight; the gated KL enters GRPO as an additive token-level advantage. Instantiated for scientific reasoning with RL-trained capability experts, VG-OPD attains the best overall performance on seven benchmarks for 4B and 8B students, ranking first on five at both scales, with the largest gains on knowledge-intensive scientific reasoning tasks. Further analysis shows that the gains come from localizing verified supervision rather than from adding teachers or distillation loss: misplacing the same supervision budget is the single most damaging change, and indiscriminate distillation drags RL below its own floor where gated distillation lifts it.

发表机构

  • Fudan University(复旦大学)
  • Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑