推理模型无需思考即可有效
Reasoning Models Can Be Effective Without Thinking
- University of California, Berkeley(加州大学伯克利分校)
- Allen Institute for AI(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究发现,通过简单提示绕过显式思考过程(NoThinking)在多个推理任务中优于传统Thinking方法,尤其在低预算场景下,且并行扩展的NoThinking方法可在相似延迟下超越基线或与高延迟Thinking方法性能相当。
AI中文摘要:
近期大型语言模型(LLMs)显著提升了推理能力,主要通过在生成过程中加入显式、冗长的思考过程实现。本文质疑这种显式思考是否必要。使用最先进的DeepSeek-R1-Distill-Qwen模型,我们发现通过简单提示绕过思考过程(记为NoThinking)的效果出人意料地好。在控制token数量时,NoThinking在七个具有挑战性的推理数据集(包括数学解题、形式化定理证明和编码)上均优于Thinking,尤其在低预算设置中,例如在ACM 23数据集上使用700个token时,NoThinking得分为51.3,而Thinking仅为28.9。值得注意的是,随着k值增加,NoThinking的pass@k性能更具竞争力。基于这一观察,我们证明了一种并行扩展方法——使用NoThinking独立生成N个输出并进行聚合——非常有效。对于聚合,我们在可用时使用任务特定验证器,或应用简单的best-of-N策略(如基于置信度的选择)。我们的方法在相似延迟下优于一系列使用Thinking的基线,且与延迟显著更长(高达9倍)的Thinking方法性能相当。总之,我们的研究鼓励重新思考冗长思考过程的必要性,同时为在低预算设置或低延迟下通过并行扩展实现强推理性能建立了竞争性参考。
英文摘要:
Recent LLMs have significantly improved reasoning capabilities, primarily by including an explicit, lengthy Thinking process as part of generation. In this paper, we question whether this explicit thinking is necessary. Using the state-of-the-art DeepSeek-R1-Distill-Qwen, we find that bypassing the thinking process via simple prompting, denoted as NoThinking, can be surprisingly effective. When controlling for the number of tokens, NoThinking outperforms Thinking across a diverse set of seven challenging reasoning datasets--including mathematical problem solving, formal theorem proving, and coding--especially in low-budget settings, e.g., 51.3 vs. 28.9 on ACM 23 with 700 tokens. Notably, the performance of NoThinking becomes more competitive with pass@k as k increases. Building on this observation, we demonstrate that a parallel scaling approach that uses NoThinking to generate N outputs independently and aggregates them is highly effective. For aggregation, we use task-specific verifiers when available, or we apply simple best-of-N strategies such as confidence-based selection. Our method outperforms a range of baselines with similar latency using Thinking, and is comparable to Thinking with significantly longer latency (up to 9x). Together, our research encourages a reconsideration of the necessity of lengthy thinking processes, while also establishing a competitive reference for achieving strong reasoning performance in low-budget settings or at low latency using parallel scaling.