自适应FastOPD:用于高效在线策略蒸馏的感知进展的展开范围扩展
Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
该研究针对在线策略蒸馏的高计算成本问题,提出自适应FastOPD策略,仅在学习平台且范围充分利用时扩展展开范围,使训练时间降49.1%-71.2%且性能最优。
中文摘要 AI 辅助
在线策略蒸馏(On-policy distillation, OPD)会沿着学生生成的轨迹提供密集的教师监督,但其在线展开过程会产生大量计算成本,尤其是当少数长响应延迟批次完成时。现有加速方法通常使用固定预算或绝对的师生一致性阈值来控制展开长度,这可能无法反映不同模型和训练阶段的学习进展。我们提出了自适应FastOPD,一种感知进展的策略,仅当在当前边界区域附近的学习已进入平台期且当前范围得到充分利用时,才扩展展开范围。前者通过相对于每次进入新范围时的师生值测量的四个师生信号确定,这使得扩展能够响应特定阶段的进展,而非预定义的步骤间隔或原始一致性信号的绝对阈值;后者则防止少数长响应触发展开成本的增加。在两对师生模型上,自适应FastOPD取得了最高的平均性能,同时相比OPD 15K将训练时间减少了49.1%至71.2%,并且在一系列超参数设置下保持鲁棒性。
英文摘要
On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion. Existing acceleration methods typically control rollout length using fixed budgets or absolute teacher--student agreement thresholds, which may not reflect learning progress across different models and training stages. We propose Adaptive FastOPD, a progress-aware strategy that expands the rollout horizon only when learning near the current boundary region has plateaued and the current horizon is sufficiently utilized. The former is determined from four teacher--student signals measured relative to their values upon entering each horizon, making expansion responsive to stage-specific progress rather than a predefined step interval or an absolute threshold on the raw agreement signals, while the latter prevents a small number of long responses from triggering increases in rollout cost. Across two teacher--student pairs, Adaptive FastOPD achieves the highest average performance while reducing training time by 49.1--71.2\% relative to OPD 15K, and remains robust across a range of hyperparameter settings.