更好的监督就在附近:邻域在策略自蒸馏
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
浏览论文内容
中文总结 AI 辅助
邻域在策略自蒸馏(N-OPSD)通过局部参数扰动构建专家池,在数学推理任务上利用参考对齐修正提升学生模型性能,在多个基准上超越标准OPSD。
中文摘要 AI 辅助
在策略自蒸馏(OPSD)使用一个特权教师来训练数学推理模型,该教师能看到参考解并监督学生采样的前缀。标准OPSD在每个状态使用一个固定的参数设置,但附近的设置可能提供额外的监督。我们发现,在相同参考上下文下,局部参数扰动揭示了互补的参考对齐修正。不同的专家在参考的不同位置提供这些修正。他们的池比未扰动的特权教师覆盖了更多这样的位置。我们引入邻域在策略自蒸馏(N-OPSD)来将这些修正转化为学生访问状态下的监督。离线时,贪心选择通过奖励每个位置上过滤后的参考令牌增益(超出池当前最佳)来构建一个紧凑的冻结专家池。最高峰专家不必提供最佳训练目标。因此,在线路由将锚定方向与其支持水平分开。MaxPeak选择锚定令牌,分位数选择在顶部令牌匹配的专家中进行选择。学生通过继承自OPSD的裁剪前向KL目标学习所选专家的完整下一个令牌分布。我们在AIME 2024、AIME 2025和HMMT 2025年2月上进行评估。在每种方法的三次独立运行中,邻域OPSD在Qwen3-1.7B、4B和8B上分别将三个基准的Average@12比OPSD提高了2.75、1.67和1.94分。学生前缀续写支持使用超出用于选择的参考轨迹的池。匹配的消融支持过滤后的参考令牌增益作为选择标准。考虑池内的重叠和按状态路由进一步提高了学生准确性。推理仅使用蒸馏后的学生模型。
英文摘要
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Meituan(美团)
- Peking University(北京大学)
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。