arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MetaOPD:面向策略内蒸馏的元学习 token 加权方法

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

Zipeng Wang, Xinpeng Dong, Yuefan Wang, Pingchen Lu, Xian Wei, Kun Kuang, Fei Wu, Zhongxiang Dai, Min Zhang

arXiv 2610.11989首次发表:更新:

发表机构

East China Normal University; Zhejiang University; The Chinese University of Hong Kong, Shenzhen(华东师范大学; 浙江大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MetaOPD 是一种双层优化框架,联合学习学生模型与轻量级 token 加权网络,在多数学推理及域外数据集上,较 OPD 显著提升了不同规模学生模型的 Avg@8、Pass@8 指标。

AI 中文摘要

策略内蒸馏(On-policy distillation, OPD)利用学生模型自身生成的响应,结合 token 级别的教师监督信号来训练学生模型。然而,均匀加权的方式忽略了不同 token 的学习价值差异,而现有的加权方法依赖于从预测信号到 token 权重的预定义映射,这些映射并非从学生模型更新的有效性中学习得到,限制了其适应不断变化的学习需求的能力。本文提出 MetaOPD,这是一种双层优化框架,可联合学习学生模型和轻量级 token 加权网络:内层目标通过加权 OPD 更新学生模型,外层目标在虚拟学生更新后,利用参考解的验证损失来优化加权网络。通过该更新过程的微分,将加权决策与其对更新后性能的影响关联起来,使预测信号到 token 权重的映射能随学生模型共同演化。在六个数学推理数据集和三个域外数据集上开展实验,覆盖两种学生模型规模和七个基准方法,结果表明 MetaOPD 有效:对于 0.6B 规模的学生模型,其 Avg@8、Pass@8 指标较 OPD 分别提升 1.99、5.97 个百分点;对于 1.7B 规模的学生模型,上述指标分别提升 2.25、6.41 个百分点。

英文摘要

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network. The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update. Differentiating through this update connects weighting decisions to their effects on post-update performance, allowing the mapping from prediction signals to token weights to evolve alongside the student. Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑