思维链(CoT)何时有助、何时有害:大语言模型推理中序列深度瓶颈的实证研究
When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning
- Vistula University(维斯图拉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究通过实证发现,思维链(CoT)是带宽旁路而非通用推理增强器,在高深度推理任务中可提升大语言模型准确率,在浅层任务中冗余,其效果与任务序列深度、模型规模相关。
AI中文摘要:
人们普遍认为思维链(CoT)提示能普遍提升大语言模型(LLM)的推理能力。我们基于H_dp带宽界(Chen等人,2024)的概念框架对此展开研究:尽管该形式界仅在渐近情形(提示长度极大时)有效,但它明确了一个真实的架构瓶颈——超出Transformer单次处理容量的序列计算必须外部化,而这正是CoT所做的。我们的核心发现是基准内的序列深度梯度:单次处理(无CoT)的准确率随单项目序列深度单调下降,而CoT则近似与深度无关。我们在实用上下文长度下,于三个指令调优模型(Qwen-2.5-7B、Qwen-2.5-32B、Llama-3.1-8B)和五个标准NLP基准中测量了CoT的效果。在高深度P-完全任务(GSM8K、MATH)上,CoT在所有模型中提供了54至68个百分点(pp)的恢复增益;在浅层TC^0任务(MMLU、ARC)上,CoT在结构上是冗余的(差值在0.0至+4.6 pp之间,无显著负面影响)——尽管无CoT基线值较高(ARC上达95%)可能反映了数据污染,因此该零结果并非纯粹的架构测试。中间类L任务(HumanEval)呈现出依赖模型规模的转变:32B模型增益+23.2 pp,8B模型增益+9.1 pp,7B模型增益-28.7 pp。跨基准的深度-恢复相关性为斯皮尔曼rho=0.661(p=0.007,n=15);经Bonferroni校正后,15个基准级McNemar检验中有9个具有统计学意义。本研究在OSF上预注册,结果表明CoT并非通用的推理增强器,而是充当带宽旁路:它有助于超出单次处理容量的序列计算,对已适配单次处理的任务则是冗余的。
英文摘要:
It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically large prompt lengths), it identifies a real architectural bottleneck -- serial computation exceeding a transformer's single-pass capacity must be externalised, which is what CoT does. Our central finding is a within-benchmark serial-depth gradient: single-pass (no-CoT) accuracy degrades monotonically with per-item serial depth, while CoT is approximately depth-invariant. We measure CoT effects across three instruction-tuned models (Qwen-2.5-7B/32B, Llama-3.1-8B) and five standard NLP benchmarks at practical context lengths. On high-depth P-complete tasks (GSM8K, MATH), CoT gives a +54 to +68 pp recovery gap across all models. On shallow TC^0 tasks (MMLU, ARC), CoT is structurally redundant (Delta in [0.0, +4.6] pp, no significant negative effect) -- though high no-CoT baselines (up to 95% on ARC) may reflect contamination, so this null is not a clean architectural test. The intermediate class L (HumanEval) shows a model-size-dependent transition: +23.2 pp (32B), +9.1 pp (8B), -28.7 pp (7B). The cross-benchmark depth-recovery correlation is Spearman rho = 0.661 (p = 0.007, n = 15); 9 of 15 benchmark-level McNemar tests are significant after Bonferroni correction. Pre-registered on OSF, our results indicate that CoT is not a universal reasoning enhancer but acts as a bandwidth bypass: it helps serial computation that strains single-pass capacity and is redundant for tasks that already fit.