发表机构
Tencent Hy; HKUST(GZ); HKUST(腾讯Hy; 香港科技大学(广州); 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Flash-OPD通过自适应轨迹级边界验证替代固定地平线控制,在在线策略蒸馏中实现2.2-7.5倍加速且不损失精度。
AI 中文摘要
在线策略蒸馏(OPD)在学生生成的轨迹上提供密集的教师监督,但生成和评估长轨迹会带来大量的训练成本。现有的加速方法通过开环轨迹调度或闭环地平线适应来降低这一成本。然而,不同轨迹之间的监督兼容性可能差异很大,使得单一的轨迹地平线难以匹配其异质的可靠长度:过短的地平线可能截断有用的监督,而过长的地平线则在可靠区域之外浪费计算。我们的关键洞察是,轨迹特定的可靠性边界不必在生成之前预测。通过将可靠性视为累积的低教师-学生兼容性事件的首次通过,该边界在采样前本质上是未知的,但是否已达到该边界可以从观察到的前缀中精确确定。基于这一洞察,我们提出了Flash-OPD,它将轨迹地平线控制转变为自适应的轨迹级边界验证。Flash-OPD将缓存生成与教师验证交错进行,并根据观察到的兼容性事件独立停止每条轨迹。为了减少验证开销,最近的事件率仅用于安排下一个验证点,而实际的停止决策始终依赖于精确的累积计数。这种分离防止了估计误差导致过早终止,同时能够在生成过程中进行高效验证。在多种数据集和教师-学生设置上的广泛实验表明,Flash-OPD相比标准OPD实现了2.2倍至7.5倍的加速,同时保持或提高了准确性。
英文摘要
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose *Flash-OPD*, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. *Flash-OPD* interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that *Flash-OPD* achieves $2.2\times$--$7.5\times$ speedups over standard OPD while maintaining or improving accuracy.
Comments21 pages, 9 figures