arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39884cs.CLcs.AIcs.LG

OPSRD:基于策略的自我角色蒸馏

OPSRD: On-Policy Self-Role Distillation

  • Zhejiang University(浙江大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang

AI总结:

提出OPSRD方法,利用固定专家角色进行无参考解决方案的基于策略的自我蒸馏,通过教师加权正向KL目标优化,在竞赛数学基准上显著提升模型准确率。

AI中文摘要:

角色提示通过专家身份激发大型语言模型的专门行为,为在艰巨任务上引导推理提供了一种轻量级方法。然而,当采样解决方案仍然不正确时,评估或蒸馏完整的角色提示答案可能会错过有用的下一词偏好。转移这些偏好还需要一个能够触及学生模型很少预测的备选方案的目标函数。我们提出了OPSRD,该方法使用固定的专家角色作为特权教学上下文,在没有参考解决方案的情况下进行基于策略的自我蒸馏。一个无角色的学生模型生成轨迹,而同一基础模型的冻结实例在其确切前缀上提供角色条件下的分布,从而暴露采样延续之外的备选方案。教师加权的正向KL目标针对学生模型低估的备选方案,并通过裁剪限制单个词汇的贡献。监督仅限于学生位置中熵最高的半数,将学习集中在预测不确定的位置。在三个竞赛数学基准上使用Qwen3-1.7B、4B和8B进行的实验表明,与推理时无角色提示的基础模型相比有改进。在评估的三种散度中,正向KL在所有规模下均实现了最高的宏平均准确率。代码可在该https URL获取。

英文摘要:

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.

补充信息

↑