arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28582cs.LG

β-OPSD:基于策略优化推导,采用自蒸馏训练

$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein, Furong Huang

首次发表
浏览论文内容

中文总结 AI 辅助

β-OPSD是将普通OPSD扩展为含可控正则化参数β的策略优化通用形式,通过蒸馏近似策略优化解,在数学推理基准上性能与稳定性均优于普通OPSD。

中文摘要 AI 辅助

在线策略自蒸馏(OPSD)是提升推理类语言模型的有前景方法,但实际应用中稳定性不足,要可靠运行常需大量工程投入。我们明确了这一难题的结构性根源:普通OPSD是更广泛策略优化族中β=1的成员,其中β用于加权将学生模型锚定到参考策略的KL散度惩罚项。这种等价关系使β从固定为1的隐式值变为可控正则化参数,得到更通用的公式,可权衡与参考策略的贴近度及特权教师指导。我们提出β-OPSD,将其最优策略推导为参考策略与特权教师的几何插值。不过,直接用强化学习优化该目标会成本高、方差大,因此我们将其闭式解转化为蒸馏目标,每个β值对应参考到教师路径上的一个目标,通过混合二者的token级logits高效实现,用低成本蒸馏近似高成本策略优化的解。同时,回报式信用分配进一步使token更新与序列级目标对齐,且保留OPSD的简洁性。在数学推理基准上的实验显示,β-OPSD始终优于普通OPSD,提升了优化稳定性和下游推理性能。我们的结果提供了从自蒸馏到策略优化再返回的原理性路径,未牺牲OPSD的实用效率。

英文摘要

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the $β=1$ member of a broader policy-optimization family, where $β$ weights the KL penalty anchoring the student to a reference policy. This equivalence turns $β$ from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce $β$-OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of $β$ selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that $β$-OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.

发表机构

  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

↑