REVO:通过方差引导重用的高效离策略蒸馏
REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse
浏览论文内容
中文总结 AI 辅助
REVO提出一种离策略蒸馏框架,通过重用学生轨迹并利用学生-教师对数概率比方差指导多步更新,在仅用50次轨迹迭代时达到或超过200次迭代的同策略蒸馏基线。
中文摘要 AI 辅助
同策略蒸馏(OPD)使用教师模型在基于学生生成轨迹上的密集词级监督来训练语言模型。然而,其对频繁更新的学生轨迹的依赖往往导致巨大的生成成本。我们提出REVO,一种离策略蒸馏框架,通过重用每条学生轨迹进行多步学习者更新来提高轨迹效率。REVO通过稳定的前缀加权和从当前学生模型进行一步重采样来解决前缀级和当前词策略不匹配问题,从而无需重新生成完整轨迹即可进行重复更新。为了在重用的轨迹中优先考虑信息丰富的词位置,REVO使用学生-教师对数概率比的方差来量化剩余的词级学习信号并指导重复优化。在多种学生-教师规模下,仅使用50次轨迹迭代的REVO在域内和跨域推理基准上达到或超过了训练200次迭代的OPD基线。
英文摘要
On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajectories. To prioritize informative token positions within reused rollouts, REVO uses the variance of the student-teacher log-probability ratio to quantify the remaining token-level learning signal and guide repeated optimization. Across multiple student-teacher scales, REVO with only 50 rollout iterations matches or exceeds OPD baselines trained for 200 iterations on both in-domain and cross-domain reasoning benchmarks.
发表机构
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
- Brigham Young University(杨百翰大学)
机构由 AI 辅助整理,请以论文原文为准。