arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DemoPSD: 分歧调节的策略自蒸馏

DemoPSD: Disagreement-Modulated Policy Self-Distillation

Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song

arXiv 2607.02502首次发表:更新:

发表机构

City University of Hong Kong; Tsinghua University; Shenzhen University of Advanced Technology; Chinese University of Hong Kong, Shenzhen(香港城市大学; 清华大学; 深圳理工大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DemoPSD框架,通过选择性采纳教师指导,平衡学生从教师学习和保持自身推理能力,解决特权信息泄露和探索抑制问题,在科学推理任务上优于现有方法。

AI 中文摘要

同策略自蒸馏(OPSD)已成为训练大型语言模型(LLMs)进行推理的实用方法,其中单个模型同时扮演教师和学生角色,但具有不同的信息访问级别。然而,最近的研究发现,教师基于特权信息的密集令牌级监督可能导致对领域内模式的过拟合、抑制探索并损害跨领域泛化,同时还引入了一个更根本的问题:*特权信息泄露*,即学生编码了在测试时不可用的依赖于答案的捷径。我们引入了**DemoPSD**,一种通过*选择性采纳教师指导*思想解决这些问题的新框架。DemoPSD不拟合完整的教师分布,而是将学生引导向一个*反向KL重心目标*,即教师和学生分布的加权几何组合,自然地在向教师学习和保持学生自身推理能力之间取得平衡。我们测量它们分布之间的差异,并利用这种差异自适应地控制每个令牌位置的混合。我们可证明地表明DemoPSD实现了**(1)** *泄露衰减*,即有效缓解特权信息泄露;以及**(2)** *探索保持*,即在密集令牌级蒸馏下保持探索能力。在四个科学领域的SciKnowEval上的大量实验表明,DemoPSD在保持更高训练熵并稳健泛化到分布外GPQA基准的同时,优于GRPO和SDPO。

英文摘要

On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access. However, recent studies have found that the teacher's dense token-level supervision, conditioned on privileged information, can lead to overfitting to in-domain patterns, suppress exploration, and hurt cross-domain generalization, while also introducing a more fundamental issue: *privileged information leakage*, where the student encodes answer-dependent shortcuts that are unavailable at test time. We introduce **DemoPSD**, a novel framework that resolves such problems through the idea of *selective adoption of teacher guidance*. Instead of fitting the full teacher distribution, DemoPSD steers the student toward a *reverse-KL barycenter target*, a weighted geometric combination of the teacher and student distributions, that naturally balances learning from the teacher with preserving the student's own reasoning capacity. We measure the difference between their distributions and use such a discrepancy to adaptively control the blending at each token position. We provably show that DemoPSD achieves **(1)** *leakage attenuation*, i.e., effective mitigation of privileged information leakage; and **(2)** *exploration preservation*, i.e., preservation of exploration capacity under dense token-level distillation. Extensive experiments on SciKnowEval across four scientific fields show that DemoPSD outperforms both GRPO and SDPO while maintaining higher training entropy and robustly generalizing to out-of-distribution GPQA benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑