CAT:面向大型推理模型高效推理的置信度自适应思考
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
- Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China(电子科技大学智能协同计算实验室)
- Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province(四川省泛在智能与可信服务重点实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出置信度自适应思考(CAT)框架,利用模型内在自确信信号作为置信度,通过偏好优化自主调节推理长度,在保持准确率的同时压缩简单问题的推理开销。
AI中文摘要:
大型推理模型(LRMs)通过利用长思维链(CoT)轨迹在复杂任务上取得了显著成功,但它们在简单查询上经常表现出过度思考,导致显著的令牌开销和推理效率降低。然而,现有的压缩方法主要应用统一的长度缩减或依赖粗粒度的难度估计,往往导致在困难问题上的性能下降。为了解决这一局限性,我们提出了置信度自适应思考(CAT)框架,该框架将模型内在的自确信信号作为置信度纳入偏好优化过程,从而根据问题难度自主调节推理长度。实验结果表明,CAT在不同基础模型的多个基准测试中,在推理准确性上始终优于最先进的基线方法。我们的工作使LRMs能够有效压缩自信响应,同时仔细考虑不确定的响应,为实际工业场景中平衡准确性和延迟提供了一种潜在的稳健解决方案。
英文摘要:
Large Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet they frequently exhibit overthinking on simple queries, resulting in significant token overhead and reduced inference efficiency. However, existing compression methods predominantly apply uniform length reduction or rely on coarse-grained difficulty estimation, often leading to performance degradation on difficult problems. To address this limitation, we propose Confidence-Adaptive Thinking (CAT), a framework that incorporates the model's intrinsic self-certainty signals as confidence into the preference optimization process, which autonomously modulates reasoning lengths based on problem difficulty. Experimental results show that CAT consistently outperforms state-of-the-art baselines on reasoning accuracy across multiple benchmarks on different base models. Our work enables LRMs to effectively compress confident responses while deliberating on uncertain ones, offering a potentially robust solution for balancing accuracy and latency in practical industrial scenarios.