arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11266cs.AI

有效≠必要:诊断思维链中的潜在低效率

Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought

Daeyeop Lee, Hwanjo Yu

首次发表
浏览论文内容

中文总结 AI 辅助

研究思维链提示中存在的有效但低效推理问题,提出基于信息论的CAID度量,应用于PACE策略,能在保持准确率时减少令牌消耗,成功从推理链中去除冗余信息。

中文摘要 AI 辅助

思维链(CoT)提示显著提升了大语言模型(LLMs)的推理能力,但由于过度推理会产生大量计算成本。现有推理步骤评估器无法惩罚有效但低效的推理步骤。为此引入RIV-GSM8K诊断基准。实验表明现有评估器难以区分低效与必要推理。提出基于信息论的无训练度量CAID识别低效用步骤,应用于PACE策略,实验证明其能在保持准确率的同时减少令牌消耗。

英文摘要

Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps. While existing reasoning step evaluators effectively detect logical fallacies and factual errors, our analysis reveals a critical blind spot: they fail to penalize valid but inefficient reasoning steps that inflate token usage without contributing to the solution. To systematically diagnose this limitation, we introduce RIV-GSM8K, a diagnostic benchmark injected with five distinct types of inefficiencies, including circular reasoning and excessive decomposition. Diagnostic experiments reveal that state-of-the-art evaluators struggle to distinguish these inefficiencies from necessary reasoning. To address this gap, we propose CAID (Context-Aware Information Density), a training-free metric grounded in information theory that identifies low-utility steps. To validate the metric's practical utility, we apply it within PACE, a post-hoc compression strategy. Additional control experiments show that the gains of PACE are not explained by trivial pruning: compared with random step removal and PRM-based compression baselines, it preserves accuracy at substantially higher compression rates. Empirical results on GSM8K, StrategyQA, and ARC-Challenge demonstrate that PACE reduces token consumption by 31-53% while maintaining accuracy, confirming that CAID successfully distills informational froth from reasoning chains without compromising deductive validity.

发表机构

  • KT Corporation(KT公司)
  • Pohang University of Science and Technology(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑