arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.00578cs.AIcs.CL

草稿式推理:在长链式推理LLM中学习高效推理

Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs

  • Zhejiang University(浙江大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Jie Cao, Tianwei Lin, Zhenxuan Fan, Bo Yuan, Ziyuan Zhao, Rolan Yan, Wenqiao Zhang, Siliang Tang

更新

AI总结:

Draft-Thinking通过草稿式推理结构和渐进式课程学习,在长链式推理LLM中实现高效推理,显著降低推理预算并保持性能

AI中文摘要:

长链式推理(CoT)已成为增强大推理模型(LRMs)推理能力的主要范式;然而,性能提升往往伴随着推理预算的显著增加。最近的研究表明,现有CoT范式倾向于引发系统性过度思考,将推理能力与推理成本不必要地耦合在一起。大多数先前方法通过事后技术如令牌压缩、截断或长度惩罚来减少令牌使用,但并未明确解决推理的核心机制。我们提出了Draft-Thinking,它引导模型首先学习一种简洁的草稿式推理结构,仅保留关键推理步骤。通过渐进式课程学习,模型在能力提升时稳定地内化这种高效的推理模式。此外,Draft-Thinking引入了自适应提示,使推理深度成为灵活、可模型选择的行为。广泛的实验表明,Draft-Thinking在大幅减少推理预算的同时,能够很大程度地保持推理性能;例如,在MATH500上,它在仅导致2.6%性能下降的情况下,实现了82.6%的推理预算减少。

英文摘要:

Long chain-of-thought~(CoT) has become a dominant paradigm for enhancing the reasoning capability of large reasoning models~(LRMs); however, the performance gains often come with a substantial increase in reasoning budget. Recent studies show that existing CoT paradigms tend to induce systematic overthinking, unnecessarily coupling reasoning capability with reasoning cost. Most prior approaches reduce token usage through post hoc techniques such as token compression, truncation, or length penalties, without explicitly addressing the core mechanisms of reasoning. We propose \textbf{Draft-Thinking}, which guides models to first learn a concise \textit{draft-style} reasoning structure that retains only the critical reasoning steps. Through a \textit{progressive curriculum learning}, the model stably internalizes this efficient reasoning pattern as its capability scales. Moreover, Draft-Thinking introduces adaptive prompting, which elevates reasoning depth to a flexible, model-selectable behavior. Extensive experiments demonstrate that Draft-Thinking substantially reduces reasoning budget while largely preserving reasoning performance; for example, on MATH500, it achieves an 82.6\% reduction in reasoning budget at the cost of only a 2.6\% performance drop.

↑