arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DIAL-OPD:在策略蒸馏中从更少的 token 中学习更多

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

Anhao Zhao, Haoran Xin, Junlong Tong, Yingqi Fan, Xuan Lu, Ping Nie, Wenjie Li, Xiaoyu Shen

arXiv 2610.11659首次发表:更新:

发表机构

Eastern Institute of Technology; The Hong Kong Polytechnic University; HKUST (GZ); Shanghai Jiaotong University; University of Hong Kong; University of Waterloo(宁波东方理工大学; 香港理工大学; 香港科技大学(广州); 上海交通大学; 香港大学; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DIAL-OPD 是一种 token 选择方法,通过加权教师与学生概率的对数均值优化策略蒸馏,在数学推理基准中仅保留 40% token 就优于基线,性能超越全 token 策略蒸馏,还能以更小规模教师模型达到更强效果。

AI 中文摘要

在策略蒸馏(On-Policy Distillation, OPD)中,教师模型会以 token 级别的信号监督学生模型生成的轨迹。其采样 token 的变体可避免全词汇概率的计算成本。然而我们发现,使用更少的 token 进行训练的效果优于全 token 的 OPD,这挑战了“更多监督会提升学习效果”的直觉,因此需要通过学习价值来选择 token。现有的基于分歧的准则忽略了概率尺度:两个模型都分配到可忽略概率的 token(称为低-低 token)可能会获得较大的对数比率奖励,进而阻碍学习。我们提出 DIAL-OPD,这是一种 token 选择方法,通过用教师模型与学生模型概率的对数均值对奖励幅度进行加权,从而连接对数概率空间与概率空间。参数 β 控制该加权过程,保留得分最高的 token。在 4 组教师-学生模型对和 7 个数学推理基准测试中,我们将 DIAL-OPD 与 9 种基线方法进行比较:仅保留 40% 的 token 时,它的性能优于 Vanilla OPD 及其全 token 变体,与 Vanilla OPD 相比平均准确率提升达 5.25 个百分点,且将 AIME25 的 Pass@16 从 13.33% 提升至 26.67%;在相同保留率下,它还比最强的 token 选择基线方法实现了最高 18% 的相对平均准确率提升。使用 4B 规模的教师模型时,DIAL-OPD 在两种学生模型规模下均超越了使用 8B 规模教师模型的最强全 token 基线方法,表明有效的监督分配可胜过教师模型的规模扩展。进一步分析显示,适中的 β 可在抑制低-低 token 与保留有用分歧之间取得平衡;token 级别的证据表明,DIAL-OPD 会过滤掉推理价值有限的高奖励 token,同时保留对推理正确性至关重要的监督信号。

英文摘要

On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑