arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OS-Pruner:通过最优停止修剪推理模型的思维链

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

Mohammed Ehab, Aymane El Gadarri, Vivek F. Farias, Adam Jozefiak, Ciamac C. Moallemi

arXiv 2607.11089首次发表:更新:

发表机构

MIT; Columbia Business School(麻省理工学院; 哥伦比亚商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型思维链推理存在计算过度问题,提出轻量级框架OS-Pruner,将思维链修剪转化为最优停止问题,通过优化效用动态评估终止点,在多基准和模型上减少生成长度且精度牺牲小。

AI 中文摘要

大语言模型通过思维链提示在复杂推理任务中取得了显著成功。然而,这些模型常出现“计算过度思考”,产生冗余推理步骤,增加延迟和成本却未提高准确性。近期研究表明思维链轨迹可大幅修剪,但现有方法存在依赖固定思维预算、启发式过滤、次优早期分类退出或昂贵再训练等问题。本文引入OS-Pruner,一个轻量级插件框架,将思维链修剪表述为最优停止问题。给定推理前缀,它通过优化权衡最终答案准确性和生成长度的显式效用,学习进一步推理是否值得其令牌成本。该新颖表述使模型能动态评估推理链的充分终止点。OS-Pruner在训练和推理时都很轻量级,能让用户对推理努力与准确性权衡进行细粒度控制。在各种推理基准和基础模型上,OS-Pruner在最小精度牺牲的情况下,生成长度减少了20%-60%。

英文摘要

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improving accuracy. Recent studies suggest that CoT trajectories can be significantly pruned, yet existing methods often rely on forcing a static thinking budget, heuristic filtering, sub-optimal early exit via classification, or expensive re-training. In this paper, we introduce OS-Pruner, a lightweight plug-in framework that formulates chain-of-thought pruning as an optimal stopping problem. Given a reasoning prefix, OS-Pruner learns whether further reasoning is worth its token cost by optimizing an explicit utility that trades off final-answer accuracy against generated length. Our novel formulation enables the model to dynamically assess the sufficient point of termination for a reasoning chain. OS-Pruner is designed to be lightweight during both training and inference, and to provide users with fine-grained control over the reasoning-effort vs. accuracy trade-off. On diverse reasoning benchmarks and base models, OS-Pruner achieves 20-60\% reduction in generation length with minimal accuracy sacrifice.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑