arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ResOPD:面向稀疏在线策略蒸馏的尾部残差化

ResOPD: Tail Residualization for Sparse On-Policy Distillation

Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen, Yuhang Zang

arXiv 2610.04882首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; University of Science and Technology of China; Fudan University(上海交通大学; 上海人工智能实验室; 中国科学技术大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ResOPD通过将未观察词汇聚合为粗粒度尾部事件并采样细粒度残差,在稀疏在线策略蒸馏中实现无偏且低方差的梯度估计,稳定训练并提升性能。

AI 中文摘要

在线策略蒸馏(OPD)在自生成的轨迹上训练学生模型,但沿长推理轨迹传递稠密的教师分布会造成严重的通信和内存瓶颈。因此,实际系统依赖稀疏的教师接口,通常仅传递采样令牌的得分或一个小的Top-$k$分布。然而,这种稀疏设置面临一个根本性困境:采样令牌估计器是无偏的,但存在严重的梯度方差,而直接优化Top-$k$目标则引入系统性偏差。为改善这一权衡,我们提出ResOPD(带尾部残差化的在线策略蒸馏),它在在线策略采样下提供无偏的全词汇表反向KL梯度估计,并在评估的设置中,在相同的稀疏负载下大幅降低方差。ResOPD将未观察到的词汇聚合为一个可观察的粗粒度尾部事件,计算其精确的聚合梯度,并仅对细粒度的尾部内残差进行采样,这无需额外的教师查询或前向传播。大量实验表明,ResOPD显著降低了梯度方差,稳定了在线训练动态,并在评估的设置中提升了下游性能。这些结果确立了ResOPD作为稀疏在线策略蒸馏的一种高效、即插即用的方差缩减原语。

英文摘要

On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled-token score or a small Top-$k$ distribution. However, this sparse setting faces a fundamental dilemma: sampled-token estimators are unbiased but suffer from severe gradient variance, whereas directly optimizing Top-$k$ objectives introduces systematic bias. To improve this trade-off, we propose ResOPD (On-Policy Distillation with Tail Residualization), which provides unbiased full-vocabulary reverse KL gradient estimation under on-policy sampling, with substantial variance reduction under the same sparse payload in the evaluated settings. ResOPD aggregates the unobserved vocabulary into an observable coarse tail event, computes its exact aggregate gradient, and samples only the fine-grained within-tail residual, which requires no additional teacher queries or forward passes. Extensive experiments demonstrate that ResOPD substantially reduces gradient variance, stabilizes online training dynamics, and improves downstream performance across the evaluated settings. These results establish ResOPD as an efficient, plug-and-play variance reduction primitive for sparse on-policy distillation.

Comments25 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑