arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特权上下文作为在线策略自蒸馏中的漂移

Privileged Context as Drift in On-Policy Self-Distillation

Ravenor Davion, Nick Rui

arXiv 2610.07842首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究在线策略自蒸馏中特权上下文的内容与来源对策略漂移的影响,发现内容变化比来源变化引起更大KL漂移,建议将特权上下文纳入稳定性设计。

AI 中文摘要

在线策略自蒸馏(OPSD)训练语言模型,使其匹配以特权上下文为条件的自身副本。现有研究在改变模型、数据和训练设置的同时,也改变了特权上下文的内容及其产生方式,这使得特权上下文设计的影响难以单独分离。受持续学习中减少灾难性遗忘工作的启发,我们研究了特权上下文的选择如何影响策略漂移。具体而言,我们变化两个轴:内容(演示、反馈或改写)和来源(外部、带验证器的自生成或不带验证器的自生成)。我们使用OPSD在九种组合和三个数据集上训练Qwen2.5-7B,测量目标任务准确率、先前任务保留率、与基础策略的逆向KL散度以及参数更新几何。固定来源时,改变内容所跨越的中位KL范围比固定内容并改变来源时更宽。这两个范围的比值对于每令牌KL为5.1倍,对于每序列KL为2.2倍。参数更新几何显示出相同模式:共享内容的适配器更新比共享来源的适配器更新更紧密对齐(平均余弦相似度为0.571对0.255)。对于持续学习,这些发现表明特权上下文应被视为OPSD稳定性设计的一部分,因为它与策略移动的距离和方向相关。

英文摘要

On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is $5.1\times$ for per-token KL and $2.2\times$ for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine $0.571$) than updates from adapters that share source ($0.255$). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.

Comments18 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑