arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11698cs.LGcs.AI

REOPD:面向在线策略蒸馏的可靠性自适应奖励外推

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出面向在线策略蒸馏的可靠性自适应奖励外推框架REOPD,通过token级系数实现细粒度可靠性自适应,在多场景下优于基线方法,无需额外模型即可提升训练稳定性与性能。

中文摘要 AI 辅助

在线策略蒸馏(OPD)在教师模型提供的密集 token 级监督下,利用学生模型自身的轨迹对其进行训练。ExOPD 等奖励外推方法会放大教师参考的对数似然比,以超越直接模仿的局限,但会对所有 token 应用单一全局系数 λ,这可能导致学生模型拟合隐式奖励的极端峰值,引发奖励黑客行为与训练不稳定问题,且最优 λ 会随领域变化,需进行代价高昂的搜索。本文提出 REOPD,一种面向 OPD 的可靠性自适应奖励外推框架。REOPD 将 token 级兼容性权重与批次级自适应预算相结合,得到 token 级系数 λ_{b,t}=1+γ_b q_t,在保留教师对齐性的同时,沿可靠的教师参考方向选择性外推。该框架无需验证器、奖励模型、价值模型或标准 OPD 之外的额外 rollout。实验结果显示,REOPD 在单教师数学任务上优于 G-OPD,在多教师设置下的两个领域均优于 G-OPD,在单教师代码任务上与 G-OPD 表现相当,证明其可在不同领域与教师配置间实现有效的细粒度可靠性自适应。

英文摘要

On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $λ$ to every token. This can drive the student to fit extreme peaks in the implicit reward, causing reward hacking and unstable training, and the optimal $λ$ varies across domains, requiring costly sweeps. We propose REOPD, a reliability-adaptive reward extrapolation framework for OPD. REOPD combines a token-level compatibility weight with a batch-level adaptive budget, yielding a token-wise coefficient $λ_{b,t}=1+γ_b q_t$ that preserves teacher alignment while selectively extrapolating along reliable teacher-reference directions. It requires no verifier, reward model, value model, or extra rollout beyond standard OPD. REOPD outperforms G-OPD on single-teacher mathematics and on both domains in the multi-teacher setting, while matching G-OPD on single-teacher code, demonstrating effective fine-grained reliability adaptation across domains and teacher configurations.

发表机构

  • Peking University(北京大学)
  • Southwest Jiaotong University(西南交通大学)
  • University of Science and Technology of China(中国科学技术大学)
  • School of Computer Science and Technology,University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)
  • University of Electronic Science and Technology of China(电子科技大学)
  • School of Automation Engineering,University of Electronic Science and Technology of China(电子科技大学自动化工程学院)
  • Renmin University of China(中国人民大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Autonomous Driving, Shanghai Artificial Intelligence Laboratory(上海人工智能实验室自动驾驶研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑