发表机构
Amazon; Duke University(亚马逊公司; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在策略内蒸馏场景中发现,极稀疏监督(每个推理轨迹仅1-2个token,占比0.05%)可媲美或超越全token训练,挑战了后训练需海量token的假设,为高效后训练算法提供新方向。
AI 中文摘要
大型语言模型通过有效的后训练展现出日益强大的推理能力。然而,主流后训练方法优化的是海量 token,隐含假设有效学习必须依赖大量 token。我们在策略内蒸馏(OPD)场景中重新审视这一假设,该场景天然允许在每个生成 token 处提供密集的教师监督。使用 Qwen3 系列模型,我们发现一种反直觉现象:推理可通过生成 token 的极小比例有效强化——每个推理轨迹仅需 1 或 2 个 token,对应所有 token 的 0.05%。令人惊讶的是,尽管训练目标排除了绝大多数生成 token,这种稀疏监督在多数情况下在提升推理能力方面可媲美或超越全 token 训练。该现象在涵盖不同模型规模的 9 种师生配置(针对数学推理任务)中一致存在,并在代码推理、Llama 模型及基于近端策略优化(PPO)与可验证奖励的强化学习(RLVR)中得到进一步验证。有趣的是,这种极稀疏监督可能更接近自然学习过程:无需逐词纠正每一步,仅反思少数关键推理步骤,更新已有认知并继续试错,在避免微观纠正的同时保持显著有效性。总体而言,我们的结果挑战了有效后训练必须依赖大量 token 的假设,并为理解和设计更高效的后训练算法指明了新方向。
英文摘要
Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of all tokens. Surprisingly, this sparse supervision in most cases matches or surpasses full-token training in improving reasoning ability, despite excluding the vast majority of generated tokens from the training objective. This phenomenon is consistently observed across nine teacher--student configurations spanning different model scales on mathematical reasoning tasks, and is further validated on coding reasoning, Llama models and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). Interestingly, such extremely sparse supervision may be closer to the natural learning process: rather than correcting every step word by word, one reflects on a few critical reasoning steps, updates prior understanding, and continues the trial-and-error, avoiding micro-level corrections while remaining remarkably effective. Overall, our results challenge the assumption that effective post-training must be token-intensive and point to a new direction for understanding and designing more efficient post-training algorithms.