arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.16315cs.AIcs.CL

激励双过程思维以实现高效的大语言模型推理

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
  • Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Wayne Xin Zhao, Xinyu Kong, Zhiqiang Zhang

更新

AI总结:

针对大型推理模型过度思考问题,提出ACPO强化学习框架,通过系统感知标记和难度估计实现自适应认知分配与动态系统切换,减少冗余推理并提升效率。

AI中文摘要:

大型推理模型(LRMs)在复杂推理任务上表现出色,但常常因过度思考而生成冗余内容,无论任务难度如何。受认知科学中双过程理论的启发,我们提出了自适应认知策略优化(ACPO),这是一种强化学习框架,使LRMs能够通过自适应认知分配和动态系统切换实现高效推理。ACPO包含两个关键组件:(1)引入系统感知推理标记,显式表示思维模式,从而使模型的认知过程透明化;(2)整合在线难度估计和令牌长度预算,以在强化学习过程中指导自适应系统切换和推理。为此,我们提出了一种两阶段训练策略。第一阶段从监督微调开始,对模型进行冷启动,使其能够生成带有显式思维模式的推理路径。在第二阶段,我们应用ACPO进一步增强自适应系统切换,以实现难度感知的推理。实验结果表明,ACPO有效减少了冗余推理,同时根据任务复杂度自适应调整认知分配,实现了高效的混合推理。

英文摘要:

Large reasoning models (LRMs) have demonstrated strong performance on complex reasoning tasks, but often suffer from overthinking, generating redundant content regardless of task difficulty. Inspired by the dual process theory in cognitive science, we propose Adaptive Cognition Policy Optimization (ACPO), a reinforcement learning framework that enables LRMs to achieve efficient reasoning through adaptive cognitive allocation and dynamic system switch. ACPO incorporates two key components: (1) introducing system-aware reasoning tokens to explicitly represent the thinking modes thereby making the model's cognitive process transparent, and (2) integrating online difficulty estimation and token length budget to guide adaptive system switch and reasoning during reinforcement learning. To this end, we propose a two-stage training strategy. The first stage begins with supervised fine-tuning to cold start the model, enabling it to generate reasoning paths with explicit thinking modes. In the second stage, we apply ACPO to further enhance adaptive system switch for difficulty-aware reasoning. Experimental results demonstrate that ACPO effectively reduces redundant reasoning while adaptively adjusting cognitive allocation based on task complexity, achieving efficient hybrid reasoning.

补充信息

↑