SAKI:面向在线策略蒸馏的最大耦合路由教师监督
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
SAKI通过KL约束的最大耦合路由教师监督,解决在线策略蒸馏中弱学生访问教师不对齐前缀的问题,在七个数学推理基准上提升Mean@8和Pass@8,并将展开吞吐量提高4.22倍。
中文摘要 AI 辅助
在线策略蒸馏(OPD)通过在学生自身生成的轨迹上训练学生来减少训练与测试状态之间的不匹配,但弱学生可能会访问教师不对齐的前缀,在这些前缀上监督的代表性较弱。我们引入了SAKI(基于KL约束插值的监督分配),该方法将KL约束的教师引导展开与最大耦合相结合,并复用已实现的接受/纠正事件来路由令牌级监督。被接受的位置保留采样令牌的反向KL监督,而纠正位置则直接对教师最高概率令牌进行监督。在最大耦合下,纠正概率恰好为TV(p_t, q_t),因此相同的信任域半径控制展开偏差,并对干预和专门监督频率进行上界约束。我们进一步实现了一个引擎驻留的投机验证器,该验证器在将匹配工作负载的展开吞吐量提高4.22倍的同时,保留了精确q轨迹分布和耦合语义。在七个数学推理基准上,SAKI在Mean@8和Pass@8指标上均提升了匹配的教师引导基线,适用于1.7B和0.6B的学生模型。位置控制和固定前缀分析进一步支持将纠正触发的路由作为冲突自适应的监督信号。
英文摘要
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
发表机构
- Meituan(美团)
- KTH Royal Institute of Technology(瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。