arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.13379cs.CLcs.AI

Thinkless:LLM学习何时思考

Thinkless: LLM Learns When to Think

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Gongfan Fang, Xinyin Ma, Xinchao Wang

更新

AI总结:

提出强化学习框架Thinkless,通过DeGRPO算法解耦控制标记与响应损失,使LLM能自适应选择短形式或长形式推理,在多个基准上将长链思考使用量减少50%-90%,显著提升计算效率。

AI中文摘要:

具备扩展思维链推理能力的推理语言模型,在需要复杂逻辑推理的任务上展现出了卓越的性能。然而,对所有查询都应用精细推理通常会导致巨大的计算低效,特别是当许多问题存在简单直接的解决方案时。这引发了一个开放问题:LLM能否学习何时进行思考?为了回答这个问题,我们提出了Thinkless,这是一个可学习的框架,使LLM能够根据任务复杂性和模型能力,在短形式和长形式推理之间进行自适应选择。Thinkless在强化学习范式下进行训练,并采用两个控制标记,<short>用于简洁回应,<think>用于详细推理。我们方法的核心是解耦组相对策略优化(DeGRPO)算法,该算法将混合推理的学习目标分解为两个部分:(1)控制标记损失,用于管理推理模式的选择;(2)响应损失,用于提高生成答案的准确性。这种解耦公式能够对每个目标的贡献进行细粒度控制,稳定了训练过程,并有效防止了在原始GRPO中观察到的坍塌现象。在实验中,在Minerva Algebra、MATH-500和GSM8K等多个基准测试上,Thinkless能够将长链思考的使用量减少50%至90%,显著提升了推理语言模型的效率。代码可在https://github.com/VainF/Thinkless获取。

英文摘要:

Reasoning Language Models, capable of extended chain-of-thought reasoning, have demonstrated remarkable performance on tasks requiring complex logical inference. However, applying elaborate reasoning for all queries often results in substantial computational inefficiencies, particularly when many problems admit straightforward solutions. This motivates an open question: Can LLMs learn when to think? To answer this, we propose Thinkless, a learnable framework that empowers an LLM to adaptively select between short-form and long-form reasoning, based on both task complexity and the model's ability. Thinkless is trained under a reinforcement learning paradigm and employs two control tokens, <short> for concise responses and <think> for detailed reasoning. At the core of our method is a Decoupled Group Relative Policy Optimization (DeGRPO) algorithm, which decomposes the learning objective of hybrid reasoning into two components: (1) a control token loss that governs the selection of the reasoning mode, and (2) a response loss that improves the accuracy of the generated answers. This decoupled formulation enables fine-grained control over the contributions of each objective, stabilizing training and effectively preventing collapse observed in vanilla GRPO. Empirically, on several benchmarks such as Minerva Algebra, MATH-500, and GSM8K, Thinkless is able to reduce the usage of long-chain thinking by 50% - 90%, significantly improving the efficiency of Reasoning Language Models. The code is available at https://github.com/VainF/Thinkless

↑