AI 中文总结
本研究首次审计了 CoT-Pass@k 的裁判验证步骤,发现三个 LLM 裁判均无法可靠检测推理链错误,且该指标在跨语言数学基准上的有效性存疑。
AI 中文摘要
Pass@k 衡量模型在重复采样下是否达到正确答案,但从不衡量其方式:一次幸运的猜测与合理的推理被同等对待。CoT-Pass@k 被提出以弥合这一差距,它增加了一个 LLM 作为裁判,在计数前必须评估解决方案的推理链。其价值完全依赖于一个假设:裁判能捕捉到有缺陷的推理。这一假设从未在该指标内部被检验过,也从未在英语之外被检验过,尽管该指标的主张涉及用于多种语言的模型。我们报告了对该验证步骤的首次审计,在该指标自身的协议下,对包含英语、土耳其语和葡萄牙语的五个数学基准的多语言套件进行了测试,其中两个基准为母语编写。我们通过确定性编辑破坏正确的解决方案,分别损害推理链和最终答案。我们观察到,所有三个裁判接受被破坏的推理链的频率几乎与接受干净链一样高。V4-Flash 和 Qwen3.6 仅在最终答案错误时才会严厉拒绝解决方案,并且当推理链与错误答案一致时更容易接受错误答案;该指标自身的裁判也接受了大多数错误答案。我们的研究表明,推理链与答案的一致性主导了两个较大裁判的裁决,且所有三个裁判均未能可靠地检测到所测试的推理错误。因此,Pass@k 与 CoT-Pass@k 之间的差异在较早的求解器生成上平均为 19.7 个百分点,而在当前生成上仅为 4.1 个百分点。所剩无几的差异取决于双方的令牌预算和生成模式;提高生成预算使 Pass@64 移动超过五十个百分点,而差异保持为零。最后,我们提出任何被评判的推理指标在将其数值视为关于推理的证据之前应通过的两项检查。
英文摘要
Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirely on one assumption: that the judge catches flawed reasoning. That assumption has never been tested inside the metric that depends on it, and never outside English, though the metric's claims concern models used in many languages. We report the first audit of that verification step, run under the metric's own protocol on a multilingual suite of five mathematical benchmarks in English, Turkish and Portuguese, two of them natively written. We corrupt correct solutions with deterministic edits that damage the chain and the final answer separately. We observe that all three judges accept corrupted chains almost as often as clean ones. V4-Flash and Qwen3.6 reject a solution sharply only when its final answer is wrong and accept a wrong answer more readily when the chain agrees with it; the metric's own judge accepts most wrong answers as well. Our study shows that chain-answer agreement dominates the two larger judges' verdicts and that all three fail to reliably detect the tested reasoning errors. Consequently the difference Pass@k - CoT-Pass@k averages 19.7 points on an earlier solver generation but only 4.1 on the current one. What little remains depends on the token budgets on both sides and on the generation mode; raising the generation budget moves Pass@64 by more than fifty points while the difference stays at zero. We close with two checks any judged reasoning metric should pass before its numbers are read as evidence about reasoning.
CommentsAccepted at MRL@EMNLP 2026 (Workshop on Multilingual Representation Learning). 24 pages, 13 figures, 6 tables