AI 中文总结
该研究发现,在相同 token 成本下,自改进、反思等含自我检查的语言模型方法均不敌重复采样,且模型规模增长会使自我检查方法的性能优势减弱,70 亿参数模型上的改写方法仍显著弱于基线。
AI 中文摘要
让语言模型规划、批评并改写自身答案、反思错误、从多次尝试中选最优或与自身副本辩论的方法,几乎都会让模型生成远多于单次思维链的文本。由于生成更多文本本身就能提升准确率,因此相较于单次思维链的优势并不能说明该方法的核心思路起到了帮助作用。Wang 等人(2024)曾报告,一个简单的基线方法——对同一问题重复采样并保留最常见答案,在预算相当的情况下往往表现更优,但该研究仅给出了点估计,未提供置信区间或显著性检验。我们将该比较作为设计实验重新开展:共包含 7 种方法、参数规模为 15 亿、30 亿和 70 亿的开源模型、2 个数学基准测试,每个测试含 150 个问题。我们统计了所有生成的 token,包括用于批评、反思、辩论轮次和检查的 token,并将每种方法与其自身实测成本下的重复采样进行比较。所有 36 项比较均按问题配对,采用自助法区间和多重性校正。在所有实验中,没有任何一种方法在相同成本下比重复采样表现更优;有 10 种方法表现明显更差,且所有这些方法都属于模型检查自身输出的类型,全部 18 项自我检查相关的比较结果均为负面。随着模型规模增长,两类自我检查方法的表现出现分化:停止选择的负面影响减弱——Best-of-N 的 8 个样本中仅统计最常见答案,在 15 亿参数模型上比让模型自行选择的准确率高出 8.0 和 11.3 个百分点,但在 70 亿参数模型上仅高出 2.0 和 1.3 个百分点,与零差异不再有显著区别;改写无法恢复性能——Self-Refine 和强制版 Reflexion 在 70 亿参数模型上仍比基线低 3.6 至 10.1 个百分点。已发表的 Reflexion 在最小规模模型上从未触发过自身重试,每次都判断自身正确,最终悄然变成了单次思维链。我们发布了代码、提示词、所有生成内容及验证脚本。
英文摘要
Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought. Because generating more text raises accuracy by itself, a gain over one chain of thought does not show the method's idea is what helped. Wang et al. (2024) reported that a simple baseline, sampling the same question repeatedly and keeping the most common answer, often wins once budgets are comparable, but gave point estimates with no confidence intervals or significance tests. We rerun that comparison as a designed experiment: seven methods, open models of 1.5B, 3B and 7B parameters, two mathematics benchmarks, 150 questions each. We count every generated token, including those spent on critiques, reflections, debate turns and checking, and compare each method against repeated sampling at its own measured cost. All 36 comparisons are paired by question, with bootstrap intervals and multiplicity correction. No method is reliably better than repeated sampling at equal cost anywhere. Ten are reliably worse, all of them methods where the model inspects its own output, and all 18 self-inspection comparisons are negative. The two kinds of self-inspection part company as models grow. Choosing stops hurting: taking Best-of-N's eight samples and just counting the most common answer beats letting the model pick by 8.0 and 11.3 points at 1.5B, but only 2.0 and 1.3 at 7B, no longer distinguishable from zero. Rewriting does not recover: Self-Refine and a forced Reflexion stay 3.6 to 10.1 points below baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and silently became a single chain of thought. We release code, prompts, all generations, and our verification scripts.