发表机构
University of Science and Technology of China; Ant Group(中国科学技术大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多轮智能体在线策略蒸馏成本高的问题,提出STRIDE方法,利用自适应早停和前缀缓冲区加速训练,在保持性能的同时实现数倍加速。
AI 中文摘要
在线策略蒸馏(OPD)已成为将能力从大型教师模型迁移到紧凑学生模型的标准方法。然而,其成本主要由自回归学生模型的展开(rollout)主导,在多轮智能体场景中扩展性较差。现有的加速方法根据固定的离线预算截断或重新定位监督信号,尽管教师信号可靠性在轨迹内部和轨迹之间都存在显著变化。我们在$\ au^2$-bench上的实证分析揭示了这种变化的清晰结构:信息丰富的监督集中在每轮的前缀部分,而且,对于多轮智能体训练最重要的是,教师认可度的跨轮丢失在时间上锁定于学生的第一个错误动作,而不是随轮次逐渐累积。基于这些发现,我们提出了STRIDE(停止与重启在线策略蒸馏加速),它结合了两种互补技术:自适应早停,一旦累积教师对数概率低于分布外阈值即终止展开;以及前缀缓冲区,缓存高质量前缀并在最弱的正确轮次处重启生成。这些机制共同诱导了一种数据驱动的课程,逐步将覆盖范围扩展到更晚的轮次。在$\ au^2$-bench零售数据集上,我们的方法匹配全轨迹OPD,并以$3.73\ imes$的加速超越30B教师模型,以$2.34\ imes$的加速超越基线本身,在跨领域多教师训练下保持$4.51\ imes$的加速。作为智能体设置之外的补充泛化测试,STRIDE在AIME 2025上以$5.10\ imes$的加速、在AIME 2024上以$3.08\ imes$的加速优于完整OPD;在两次评估的平均值上,两个固定预算截断基线仍低于完整OPD。
英文摘要
On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic settings. Existing acceleration methods truncate or relocate the supervision signal according to fixed, offline budgets, despite substantial variation in teacher-signal reliability both within and across trajectories. Our empirical analysis on $τ^2$-bench reveals a clear structure in this variation: informative supervision is concentrated in the prefix of each turn, and, most importantly for multi-turn agentic training, the cross-turn loss of teacher endorsement is temporally locked to the student's first erroneous action rather than accumulating gradually over turns. Building on these findings, we propose STRIDE (Stop-and-Restart on-policy Distillation acceleration), which combines two complementary techniques: adaptive early stopping, which terminates a rollout once the cumulative teacher log-probability falls below an out-of-distribution threshold, and a prefix buffer, which caches high-quality prefixes and restarts generation at the weakest correct turn. Together, these mechanisms induce a data-driven curriculum that progressively extends coverage to later turns. On $τ^2$-bench retail, our method matches full-trajectory OPD and exceeds the 30B teacher at a $3.73\times$ speedup, surpasses the baseline itself at $2.34\times$, and retains a $4.51\times$ speedup under cross-domain multi-teacher training. As a supplementary generalization test beyond the agentic setting, STRIDE outperforms full OPD on AIME 2025 at a $5.10\times$ speedup and on AIME 2024 at a $3.08\times$ speedup; averaged across the two evaluations, both fixed-budget truncation baselines remain below full OPD.
Comments14 pages, 8 figures