发表机构
The Hong Kong University of Science and Technology (Guangzhou); The University of Edinburgh; The Hong Kong University of Science and Technology(香港科技大学(广州); 爱丁堡大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TV-OPD方法,通过总变差正则化塑造有界且平滑的优势信号,解决在线策略蒸馏中监督信号高方差导致的不稳定问题,实现更优性能和更低方差。
AI 中文摘要
在线策略蒸馏(OPD)促进了大语言模型(LLMs)在后训练阶段从领域专家向学生模型的知识迁移。然而,主流OPD方法中的监督信号存在高方差和噪声,这在训练过程中通常是不稳定的。在这项工作中,我们系统地研究了真正影响性能的因素以及训练过程中不稳定性背后的基本机制。我们发现,仅保留token级优势的符号就足以实现与标准OPD相当的性能。同时,更平滑且有界的优势可以稳定训练过程而不牺牲其性能。这些发现促使我们使用总变差(TV)来塑造优势,并提出了一种鲁棒的TV正则化在线策略蒸馏(TV-OPD)方法。得益于有界且递减的优势,TV-OPD表现出稳定的训练动态和稳定的后期性能。我们进行了全面的实验,发现在各种设置下,TV-OPD在训练后期始终实现了更好的性能和更低的方差。
英文摘要
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the Total Variation (TV) and propose a robust TV regulated On-Policy Distillation (TV-OPD) method. Benefiting from the bounded and diminished advantages, TV-OPD exhibits stable training dynamics and steady late-stage performance. We conducted comprehensive experiments and found that, across various settings, TV-OPD consistently achieved better performance and lower variance in the late-stage of training.