SimpleOPD:适用于长上下文推理的、简单的与分词器无关的在线策略蒸馏
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出SimpleOPD方法,通过共享文本空间对齐令牌、引入学生参考KL损失并屏蔽特殊终止令牌优势,将长上下文模型SU-01的推理能力迁移至多类学生模型,在数学及科学基准上取得显著提升。
AI中文摘要:
在线策略蒸馏(OPD)为从更强的教师模型迁移推理能力提供了可行途径,但将其应用于长上下文推理教师与短上下文学生的场景时,会面临实际挑战,包括分词器不匹配、师生分布不匹配、响应长度爆炸以及训练不稳定。本研究通过将长上下文推理模型SU-01的证明推理能力迁移至短上下文学生模型来探究该场景。为处理分词器差异,我们在共享文本空间中执行OPD,仅对齐师生分词器下占据相同文本跨度的令牌。为缓解生成长度过长与频繁截断的问题,我们引入学生参考KL散度损失,并屏蔽</think>、<|im_end|>等特殊终止令牌的优势。该策略约束学生不过度偏离初始策略,从而缓解师生分布不匹配问题并促进长度稳定增长。在同系列与不同系列学生模型(包括Qwen3、Qwen3.5、Intern-S2、GLM-4.7、Gemma-4)上的实验显示,模型的数学推理能力均获得一致提升,尤其是自然语言数学证明任务。值得注意的是,Intern-S2-Preview在ProofBench上提升21.2个百分点,达到55.2,超过Gemini-2.5-Pro;同时在HLE、HiPhO等科学基准上也有提升,表明OPD可迁移超出数学训练领域的通用推理能力。
英文摘要:
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability. In this work, we study this setting by transferring proof-reasoning capabilities from the long-context reasoning model SU-01 to short-context student models. To handle tokenizer differences, we perform OPD in a shared text space and align only tokens that occupy identical text spans under the student and teacher tokenizers. To mitigate the problem of excessive generation length and frequent truncation, we introduce a student reference KL loss and mask the advantages of special termination tokens such as </think> and <|im_end|>. This strategy constrains the student from drifting excessively from its initial policy, thereby mitigating the teacher-student distribution mismatch problem and fostering steady length growth. Experiments on both same-family and different-family student models, including Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4, show consistent gains in mathematical reasoning, especially natural-language math proving. Notably, Intern-S2-Preview improves by 21.2 points on ProofBench, reaching 55.2 and surpassing Gemini-2.5-Pro. It also improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.