发表机构
Sungkyunkwan University; Microsoft(成均馆大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型推理模型对简单问题过度思考、对困难问题思考不足的低效问题,提出When2Think后训练框架,通过实例级难度感知控制动态分配计算,在数学基准上显著提升准确率-效率权衡。
AI 中文摘要
大型推理模型(LRMs)在复杂任务上表现出色,但存在系统性低效问题:它们常常对简单问题过度思考,而对困难问题思考不足。现有基于统一长度惩罚或刚性路由的方法会带来效率税,即以牺牲困难实例的准确性为代价来减少简单实例的计算量。我们将高效推理表述为一个实例自适应计算分配问题,并提出When2Think,一个面向混合推理的后训练框架,可根据问题难度动态分配计算资源。我们的方法引入了实例级难度感知控制(IDAC),这是一种奖励塑形机制,利用预先计算的参考统计量(准确率和令牌使用量)来调节推理深度。结合基于验证器的奖励和批量标准化优势,IDAC实现了无需评论家的稳定优化,无需学习奖励模型或在线参考模型查询。When2Think鼓励对简单实例直接作答,同时保留对困难实例的扩展推理,从而学习何时使用系统1(不思考)与系统2(思考)。数学基准上的实验展示了改进的准确率-效率权衡:在AIME24上,Pass@3提高了10.0%,同时令牌使用量相对于基础模型减少了27.9%;在AIME25上,When2Think实现了40.0%的Pass@3,优于压缩和仅路由基线。
英文摘要
Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly control whether}to reason and how much computation to allocate within reasoning. We call the resulting difficulty-dependent loss in accuracy under computation reduction the efficiency tax. We propose When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation. Its core mechanism, Instance-level Difficulty-Aware Control (IDAC), uses cached reference statistics of success and token cost to modulate a correctness-gated efficiency bonus based on generated token count. Importance sampling supports exploration of Think and NoThink, while Batch-Wise Standardization constructs standardized advantages for critic-free optimization. The framework requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates. On AIME24, When2Think improves Pass@3 by 10.0 percentage points while reducing token usage by 27.9% relative to the backbone.