OBC-Prune:大型推理模型剪枝的基于结果的校准
OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
浏览论文内容
中文总结 AI 辅助
提出OBC-Prune方法,通过基于结果的因果重要性校准剪枝数据,保留关键推理回路,在多个模型和稀疏度下提升剪枝后推理模型的准确率与效率。
中文摘要 AI 辅助
大型推理模型(LRMs)在回答问题之前会生成长链思维轨迹,从而产生显著的推理开销。剪枝可以降低这一成本,但其有效性取决于用于估计参数重要性的校准数据。近期工作使用模型自身的展开(rollouts)而非通用数据集进行校准,但将所有推理标记一视同仁,无论它们是否有助于成功推理。因此,剪枝根据统计显著性而非对正确推理的贡献来保护权重,导致错误计算背后的权重与正确计算背后的权重一样被保留。这些错误模式随后被带入剪枝后的模型,降低推理质量,导致准确率下降和推理轨迹变长。我们提出大型推理模型剪枝的基于结果的校准(OBC-Prune)以弥补这一差距。OBC首先从模型回答不一致的问题中构建难度匹配的正确与错误展开对。然后,通过基于干预的分析估计每个推理句子的因果重要性,量化移除其影响对后续预测的影响。这些因果重要性分数被转换为逐标记权重,用于重新缩放一次性剪枝方法(SparseGPT、Wanda、ALPS)使用的校准激活,而不修改底层剪枝算法。在DeepSeek-R1-Distill-Qwen 1.5B、7B和14B模型上,在40%和50%稀疏度下进行的实验表明,在MATH500、LiveCodeBench和AIME 2025上,与最先进的校准基线相比,在大多数模型规模和稀疏度水平上均取得了一致的改进。这些结果表明,保留因果重要的推理回路是比统一保留观察到的激活更有效的剪枝目标。
英文摘要
Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, pruning protects weights by statistical salience rather than by their contribution to correct reasoning, so weights behind erroneous computation survive as readily as those behind correct computation. These erroneous patterns then get carried into the pruned model, degrading reasoning quality, producing both lower accuracy and longer reasoning traces. We propose Outcome-Based Calibration for Large Reasoning Model Pruning (OBC-Prune) to close this gap. OBC first constructs difficulty-matched pairs of correct and incorrect rollouts from problems the model answers inconsistently. It then estimates the causal importance of each reasoning sentence through intervention-based analysis, quantifying how removing its influence affects subsequent predictions. These causal importance scores are converted into per-token weights that rescale the calibration activations used by one-shot pruning methods (SparseGPT, Wanda, ALPS), without modifying the underlying pruning algorithms. Experiments on DeepSeek-R1-Distill-Qwen 1.5B, 7B, and 14B models at 40\% and 50\% sparsity demonstrate consistent improvements over state-of-the-art calibration baselines across most model sizes and sparsity levels on MATH500, LiveCodeBench, and AIME 2025. These results indicate that preserving causally important reasoning circuits is a substantially more effective pruning objective than uniformly preserving observed activations.
发表机构
- VinUniversity
- Johns Hopkins University(约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。