AI 中文总结
本研究揭示推理模式会集中模型错误,使自洽性失效;在五个基准和74,944样本上,推理使错误答案重合率上升,且多数投票优于加权投票,置信度信号需在决策层面评估。
AI 中文摘要
自洽性假设,当模型不确定时,独立样本会产生分歧,因此一致性是正确性的证据。在固定权重并仅切换推理模式的情况下,跨越五个基准和74,944个样本,我们表明推理会集中模型的错误:两个独立抽取的错误答案重合的概率在所有十个数据集规模的比较中均上升(p = 0.00098),并且在将两个分支都限制在各自出错的题目上后,九个比较中有九个上升。在答案空间无界的情况下,推理将产生的不同答案数量削减至非推理计数的0.43-0.65倍;在答案空间有界的情况下,两个分支持有相同的选项集,推理反而将质量集中在该选项集上,这在固定权重下无法用任何位置先验解释。总体成本小于机制预测,因为推理还缩小了答案多样性能够决定任何结果的题目集合,在十个单元中全部如此,且缩小了2.7倍;针对可用余量进行归一化后,两个分支在领域内转换了其中四分之一。置信度加权无法恢复剩余部分。在八个模型和五个基准上的280种方法-数据集-模型组合中,没有一种在修正后超过简单多数投票;加权投票在98.5%的问题-方法对上与多数投票一致,在其余部分上56.3%的时间正确;并且信号的方向可以在固定权重内反转,答案对数概率在推理关闭时预测正确性,在推理开启时预测错误。一个学习到的六信号组合在领域外毫无收获。置信度信号应在决策上评估,而非在区分度上评估。
英文摘要
Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model's errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning count; where it is bounded, both arms hold an identical option set and reasoning concentrates mass on it instead, which no positional prior can explain at fixed weights. The aggregate cost is smaller than the mechanism predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything, in ten of ten cells and by 2.7x; normalized for available headroom, both arms convert a quarter of it in domain. Confidence weighting does not recover what is left. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction; weighted voting agrees with it on 98.5% of problem-method pairs and is right 56.3% of the time on the rest; and a signal's direction can invert within fixed weights, with answer log-probability predicting correctness when reasoning is off and error when it is on. A learned six-signal combination gains nothing out of domain. Confidence signals should be evaluated on decisions, not on discrimination.
Comments25 pages, 5 figures, 21 tables