AI 中文总结
该研究提出基于GRPO的自适应推理模型,通过三种模式选择分配测试时计算,在MATH数据集上减少41%响应长度且保留高准确率,还可迁移至其他基准。
AI 中文摘要
采用强化学习训练的推理语言模型通常在固定的token预算下运行,而非显式自适应预算,这会导致在简单问题上计算过度,在困难问题上计算不足。我们研究模型能否学会分配自身的推理 effort,具体方法是在响应的第一个token处选择三种模式之一:\textsc{NoThink}(尽快给出答案)、\textsc{Short}(简短推理)或\textsc{Long}(扩展推理)。该选择在Group Relative Policy Optimization(GRPO)中学习,无需单独的路由模块,通过一种塑形奖励(使每种模式在不同响应长度下具有价值)以及严格的每模式token上限(保持模式间的区分度)实现。在基于MATH训练的15亿参数蒸馏模型上,三种模式会出现且不会坍缩为单一选择,其中简短模式的准确率最终高于\textsc{Long}模式,表明路由模块是按问题难度排序而非随机选择。在三个随机种子的平均结果中,所得策略在保留的MATH500数据集上与基础模型的准确率接近(0.782 vs. 0.796),同时将平均响应长度从4796个token降至2811个token(减少41%)。有趣的是,该模型无需重新训练即可迁移到其他基准测试,在问题更简单的场景中节省量最大,例如在GSM8K上实现76%的token减少,且在相近响应长度下比基线模型准确率更高。简言之,我们构建了一种推理模型,可针对每个问题自适应选择推理程度。
英文摘要
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. We study whether a model can learn to allocate its own reasoning effort by choosing, as the first token of its response, one of three modes: \textsc{NoThink} (answer as quickly as possible), \textsc{Short} (brief reasoning), or \textsc{Long} (extended reasoning). The choice is learned inside Group Relative Policy Optimization (GRPO) with no separate router, through a shaped reward that makes each mode worthwhile at a different response length, together with hard per-mode token caps that keep the modes distinct. On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing to a single choice, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random. Averaged over three seeds, the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction). Interestingly, it also transfers to other benchmarks without retraining, with the largest savings where problems are easier, with for instance 76\% token reduction on GSM8K and at higher accuracy than the baselines at similar response length. In short, we build a reasoning model that adaptively chooses how much to reason for each problem.