从失调到编排:教师干预在在线策略蒸馏中的应用
From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对在线策略蒸馏中教师干预分配问题,提出MAESTRO方法,利用局部策略分歧动态调整干预时机与长度,在八个数学推理基准上提升学生模型准确率并减少训练响应长度。
中文摘要 AI 辅助
在线策略蒸馏(OPD)利用更强教师的反馈,在自身推理轨迹上训练学生模型。教师干预可以改进这些轨迹,但也会改变学生学习的数据分布。我们的对照研究表明,仅凭轨迹质量不足以作为分配教师指导的标准。更深入的干预在提升轨迹准确率方面收益递减,同时增加离策略负担。在限制轨迹长度的训练探针中,学生峰值准确率和性能保持性对干预强度的偏好不同。最佳干预深度和位置也因基准而异。这些发现促使我们提出MAESTRO,它利用局部策略分歧来联合调整教师接管时机和生成长度。其策略分歧分数结合了教师加权候选覆盖率和局部分布相似性,并在推理段落内聚合。在八个数学推理基准上,MAESTRO在0.6B和1.7B Qwen3学生模型中均取得了最高的宏平均准确率,其中1.7B学生在每个基准上均领先。MAESTRO还将平均训练响应长度相对标准OPD减少了67.3%。代码可在该https URL获取。
英文摘要
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3\% relative to standard OPD. The code is available at https://github.com/yhao-wang/MAESTRO.
发表机构
- Nanyang Technological University(南洋理工大学)
- Baidu Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。