arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越最优教师:扩展与压缩推理解流形

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

Songshuo Lu, Zhi Chen, Yaohua Tang

arXiv 2607.27770首次发表:更新:

发表机构

Moore Threads AI(摩尔线程人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出“先扩展后压缩”框架,结合RGRPO与TU-OPD等技术,构建互补教师并集并压缩,使Qwen3-1.7B学生模型在数学推理等三领域优于最强单教师,实现性能提升。

AI 中文摘要

单次强化学习(RL)运行可生成强推理器但教师能力不全:它通常仅放大部分有效解模式。我们认为,经RL训练的策略应被视为多盆推理解流形的局部探针,而非全局可靠监督者。基于此,我们提出“先扩展后压缩”框架,将教师构建与多教师策略蒸馏结合。扩展阶段采用残差组相对策略优化(RGRPO),从同一初始化训练一系列教师,每后续一轮均转向累积教师并集尚未覆盖的示例。压缩阶段采用可靠性门控教师并集在线蒸馏(TU-OPD),让学生从自身响应前缀学习:每个示例仅可靠教师参与,其采样token的OPD损失按每个示例质量加权。我们还引入共识残差分解,保留胜出教师相对于可靠同行的超额token偏好,防止教师聚合期间专业行为被抑制。数学推理、代码生成与指令遵循实验显示,所得Qwen3-1.7B学生模型在三个领域均持续优于最强单教师,相对提升分别为2.0%、8.3%、6.9%,同时保持单模型推理。这些结果确立了一个简单却强大的原则:更强学生模型并非通过选择单个更优教师获得,而是通过刻意构建并压缩互补教师并集实现。

英文摘要

A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore be viewed as local probes of a multi-basin reasoning solution manifold, rather than as globally reliable supervisors. Based on this view, we propose an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation. In the expansion stage, Residual Group Relative Policy Optimization (RGRPO) trains a sequence of teachers from a common initialization and redirects each later round toward examples not yet covered by the accumulated teacher union. In the compression stage, reliability-gated Teacher-Union On-policy Distillation (TU-OPD) lets the student learn from its own response prefixes. For each example, only reliable teachers contribute, and their sampled-token OPD losses are weighted by their per-example quality. We further introduce Consensus-Residual Decomposition, which preserves a winner teacher's excess token preferences over its reliable peers, preventing specialist behavior from being suppressed during teacher aggregation. Experiments on mathematical reasoning, code generation, and instruction following show that the resulting Qwen3-1.7B student consistently outperforms the strongest individual teacher across all three domains, yielding relative improvements of 2.0%, 8.3%, and 6.9%, respectively, while retaining single-model inference. These results establish a simple but powerful principle: stronger students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑