arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从思考模式优势中学习:基于策略的蒸馏

Learning from Think-Mode Advantage via On-Policy Distillation

Wanqi Ren, Jianxiang Wang, Danxuan Liu, Linyi Ding, Huaixiao Tou

arXiv 2609.37044首次发表:更新:

发表机构

ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ThinkOPD,一种基于策略的蒸馏方法,通过轨迹-响应散度代理和组相对奖励增益路由监督,在数学推理和代码生成中优于基线,实现教师思考模式优势向学生的有效转移。

AI 中文摘要

显式的中间推理赋予大型语言模型(LLMs)更强的问题解决模式。我们研究通过基于策略的蒸馏(OPD)从这种思考模式优势中学习。OPD保留学生生成的轨迹,并在学生访问的前缀处提供密集的、词元级别的教师目标。特权推理在蒸馏期间使用,而非在学生推理时使用。Uniform ThinkOPD是一种自然的、启用思考的OPD基线,它使固定教师基于一个共享的思考轨迹,并统一蒸馏每个兄弟学生响应。尽管其前缀是基于策略的,但该轨迹不必遵循与每个完整响应兼容的路径:即使响应达到相同结果,同一特权轨迹也可能导致不同的教师-学生差异。我们将这种交互总结为轨迹-响应散度(TRD),并引入ThinkOPD,它通过结合组相对奖励增益和基于TRD的兼容性代理,在响应级别路由监督。最终响应权重在每个滚动组内归一化。在数学推理和代码生成任务中,ThinkOPD在相同模型设置和两个跨模型教师-学生对中均优于Uniform ThinkOPD,并在受控比较中超过具有代表性的推理和自蒸馏基线。受控干预表明,结果收益和基于TRD的代理在此设置中提供互补的路由信号。启用思考的OPD为研究教师优势如何沿学生响应变得可转移提供了一个受控环境。

英文摘要

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome. We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy. Final response weights are normalized within each rollout group. Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison. Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting. Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑