发表机构
College of Information and Intelligence, Hunan Agricultural University; Yuelushan Laboratory(湖南农业大学信息与智能科学学院; 岳麓山实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以GPT-4.1为案例,将LLM自一致性的错误共识一致性分解为机械成分与偏好残差,发现难题上存在自一致性反效果,一致性为分级证据,未提出新投票方法。
AI 中文摘要
对多个大语言模型(LLM)样本进行多数投票被广泛用于提升答案准确性,但其提升效果波动不定:对于难题,甚至可能产生反效果。本文对这种失败进行了量化分析。定义了多元一致性指数Gamma,其含义为错误运行的样本中与共识一致的预期比例,以参考尺度d=(1-p)/(C-1)进行归一化,并将其分解为机械成分(仅基于每个案例的答案偏好投票所产生的结果)和偏好无法解释的残差。机械零值与难度匹配且无漏洞:每个案例根据其自身的准确率和选项偏好进行重新模拟,这些参数由该案例的其他运行估计,因此没有运行会预测自身的一致性。在GPT-4.1上的分解显示出与基准相关的方向(每个基准n=4个单元的观测排序,非显著性声明)。在多项选择题GPQA-Diamond上,每个案例的答案偏好解释了保留测试运行一致性指数的81%-93%:共享偏差主导的观点在此处言过其实,因为整个群体所依附的错误但有吸引力的选项被每个案例的偏好通道捕获(该偏好是否由共享训练偏差引起尚未确定)。在开放域AIME上,机械偏好仅解释了59%-78%(若收缩为纯噪声则为21%-29%),且存在1.56-2.80 Gamma单位的偏好无法解释残差,这一残差可被运行级别的偏好异质性参考完全吸收(1.4-2.1)。在难题上重现了自一致性反效果(分组投票差距低至-0.09,95%置信区间[-0.12,-0.07]),且最高一致性分组的准确率仅为0.42-0.83,比基线提升1.2-3.6倍:一致性是分级证据,而非认证。本文未提出新的投票方法,代码和证据已提交且可复现。
英文摘要
Agreement among repeated samples of a language model is routinely read as evidence about answer reliability, yet wrong answers can agree just as strongly as right ones. This paper asks what information wrong-consensus agreement actually contains, and answers with a quantitative decomposition. A pluralistic agreement index Gamma, normalized by the reference scale d=(1-p)/(C-1), is split into a mechanical component (agreement delivered by a per-case answer preference alone) and a preference-unexplained residual. The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only. On public GPT-4.1 per-run data, coverage phi (the mechanical/empirical ratio) shows a benchmark-associated direction: 0.81-0.93 on multiple-choice GPQA-Diamond against 0.59-0.78 on open-domain AIME, where a residual of 1.54-2.80 Gamma units survives, more than absorbed by a calibrated run-level preference-heterogeneity reference. A controlled replication under one fixed protocol (four runs per question, K=32 votes) on five open-weights checkpoints (Qwen3.5-9B/122B, Qwen3.8-27B, Gemma4-26B/31B) finds near-complete mechanical coverage in all ten cells (phi approximately 1, with a small overshoot consistent with a quantified finite-donor plug-in bias), robust to a two-run design; the largest cell (qwen3.5-122b, p=0.222) sits inside the GPT-4.1 AIME accuracy range and still saturates (phi=1.041). A cross-system contrast at comparable aggregate accuracy contrasts near-complete mechanical agreement in the open-weights models against a larger preference-unexplained residual in the frontier family. This contrast is confounded with sampling protocol by design. Agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed.
Commentsv2: Added controlled open-weights replication (five models, ten cells), finite-donor bias calibration, and within-family regime analysis