ChainPrune:评估与减少长思维链推理中的冗余
ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning
浏览论文内容
中文总结 AI 辅助
ChainPrune 是一种推理路径语义结构优化方法,通过整合推理路径为树结构、选择主导路径及结合 DPO 与监督损失,减少长思维链冗余,在保精度的同时降低了步骤长度与计算开销。
中文摘要 AI 辅助
思维链(CoT)推理通过引入显式中间推理,显著提升了大语言模型(LLM)的多步问题解决能力。然而,先进的大推理模型(LRM)常出现过度思考行为,包括推理步骤过长、存在冗余步骤以及计算开销高。现有的基于 token 长度的奖励策略旨在促进简洁输出,但常导致伪简洁性——token 数量减少却仍存在冗余推理,形成更长且结构效率更低的推理链。为解决这些局限,我们提出 ChainPrune,一种新颖的推理路径语义结构优化方法,用于高效且可控地合成自生成的高质量训练数据。我们首先将自生成的推理路径整合为基于树的结构,随后通过多准则主导路径选择过程构建偏好数据,形成浅层推理轨迹同时保留必要推理步骤。为进一步提升推理质量,我们结合基于 DPO 的偏好学习方法与监督损失,有效缓解错误奖励抑制。这一创新整合显著提升了我们推理框架的效率与有效性。综合实验结果表明,该方法在保持甚至提升准确性的同时,大幅降低了步骤长度与计算开销。
英文摘要
Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoning. However, advanced Large Reasoning Models (LRMs) often exhibit overthinking behaviors, including excessively long reasoning steps, redundant steps, and high computational overhead. Existing token-length reward strategies aim to promote concise outputs, but often result in pseudo-conciseness, where token count is reduced, yet redundant reasoning persists, leading to longer and less structurally efficient chains. To address these limitations, we propose ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data. We initially consolidate self-generated reasoning paths into a tree-based structure, followed by a multi-criteria dominant path selection process for preference data construction that formulates shallow reasoning trajectories while preserving essential reasoning steps. To further enhance the quality of reasoning, we incorporate a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression. This innovative integration significantly enhances both the efficiency and effectiveness of our reasoning framework. Comprehensive experimental results demonstrate significant reductions in step length and computational overhead, while maintaining or even enhancing accuracy.
发表机构
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。