学习何时思考:通过多阶段强化学习塑造R1风格模型的自适应推理
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Pengcheng Laboratory(鹏城实验室)
- School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大型推理模型过度思考问题,提出AutoThink多阶段强化学习框架,利用提示中省略号触发可控性,动态决定是否显式推理,在五个数学基准上实现准确率与效率的优化,可无缝集成于R1风格模型。
AI中文摘要:
大型推理模型(LRMs)擅长在生成最终答案之前产生显式的、逐步的推理序列。然而,这种详细的推理可能会带来大量的计算开销和延迟,尤其是对于简单问题。为了解决这种过度思考问题,我们探索如何使LRMs具备自适应思考能力:使它们能够根据问题复杂性动态决定是否进行显式推理。基于R1风格的蒸馏模型,我们观察到在提示中插入一个简单的省略号(“...”)可以随机触发思考或不思考模式,揭示了推理行为中潜在的可控性。利用这一特性,我们提出了AutoThink,一个多阶段强化学习(RL)框架,通过阶段性的奖励塑造逐步优化推理策略。AutoThink学习仅在必要时调用显式推理,而对于较简单的任务则默认给出简洁的响应。在五个主流数学基准上的实验表明,与最近的提示和基于RL的剪枝方法相比,AutoThink实现了良好的准确率-效率权衡。它可以无缝集成到任何R1风格模型中,包括蒸馏和进一步微调的变体。值得注意的是,AutoThink在DeepSeek-R1-Distill-Qwen-1.5B上将相对准确率提高了6.4%,同时将令牌使用量减少了52%,为LRMs建立了一种可扩展且自适应的推理范式。项目页面:https://github.com/ScienceOne-AI/AutoThink。
英文摘要:
Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities: enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy-efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4 percent while reducing token usage by 52 percent on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.