arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越教师似然:面向长上下文推理的组校准在线策略蒸馏

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou

arXiv 2608.19181首次发表:更新:

发表机构

Tsinghua University; Beijing University of Posts and Telecommunications; OpenBMB(清华大学; 北京邮电大学; OpenBMB)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长上下文推理中 OPD 的教师-验证器分歧问题,提出 GC-OPD 方法,结合验证器结果优化 OPD,在五个长上下文基准上显著提升 Qwen3 模型性能。

AI 中文摘要

在线策略蒸馏(OPD)利用更强教师模型的密集 token 级指导,在学生模型自身的响应上对其进行训练。然而在长上下文任务中,token 级教师指导可能偏向局部合理的响应,这类响应会遗漏分布在输入中的证据或违反全局任务约束。相比之下,特定任务的验证器会在响应层面评估任务完成情况,可能返回反映部分成功的分级奖励。我们在两个代表性长上下文证据聚合任务的固定响应上诊断了这种不匹配问题:随着输入范围变长,轨迹级 OPD 分数与验证器奖励的一致性逐渐降低,这表明存在教师-验证器分歧。受此观察启发,我们提出组校准在线策略蒸馏(GC-OPD),该方法在每个 rollout 组内分别对验证器奖励和轨迹级 OPD 分数进行归一化,将两者的差值作为带符号的教师-验证器分歧残差;基于相对优势的信用分配(RACA)则根据 token 的相对 OPD 优势,在各 token 间分配该轨迹级残差,同时保留原始 OPD 信号。在五个长上下文基准上,使用 GC-OPD 微调后,官方 Qwen3-4B 和 Qwen3-8B 检查点的五基准平均值分别从 29.08 提升至 40.47,从 35.12 提升至 44.65;相同设置下,普通 OPD 分别达到 39.31 和 43.56。控制 ablation 实验显示,带符号残差比额外 OPD 衍生项或直接添加组归一化验证器奖励更有效,而 RACA 进一步优于均匀 token 分配。这些结果表明,组相对残差校准可在不丢弃密集 token 级指导的情况下融入验证器结果,代码可访问此 https URL。

英文摘要

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses from two representative long-context evidence-aggregation tasks. Across longer input ranges, trajectory-level OPD scores become progressively less aligned with verifier rewards, indicating teacher-verifier disagreement. Motivated by this observation, we introduce Group-Calibrated On-Policy Distillation (GC-OPD). GC-OPD separately normalizes verifier rewards and trajectory-level OPD scores within each rollout group and uses their difference as a signed teacher-verifier disagreement residual. Relative-advantage-based credit assignment (RACA) distributes this trajectory-level residual across tokens according to their relative OPD advantages while preserving the original OPD signal. Across five long-context benchmarks, post-training with GC-OPD raises the five-benchmark averages of the official Qwen3-4B and Qwen3-8B checkpoints from 29.08 to 40.47 and from 35.12 to 44.65, respectively. Vanilla OPD reaches 39.31 and 43.56 under the same setup. Controlled ablations show that the signed residual is more effective than either an additional OPD-derived term or direct group-normalized verifier reward addition, while RACA further improves over uniform token allocation. Together, these results demonstrate that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance. Code is available at https://github.com/SolereZhang/GC-OPD.

Comments20 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑