当大语言模型达成一致时,它们是正确的吗?将自一致性和跨模型一致性作为置信信号进行审计
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
浏览论文内容
中文总结 AI 辅助
研究探讨大语言模型一致性与正确性的关系,通过大规模交叉运行研究,发现一致性是正确性的条件代理,对不饱和中层模型和计算分配较有用,对最一致的前沿模型存在过度自信问题,还公开了相关数据。
中文摘要 AI 辅助
大语言模型作为评判器在企业管道中越来越多地成为评估人工智能系统的默认方式,通常扩展为集成或“专家混合”评判小组。这些系统都有一个关键假设:一致性表明正确性。但我们发现这个假设不可靠。一致性并非准确性,模型可能因共同偏差、记忆启发式或选项位置先验而非事实达成一致。我们通过大规模交叉运行研究询问何时一致性仍是可用代理。53个运行器针对GPQA Diamond和AIME上模型层级、提示和规模的比较抽取样本。结果表明一致性是一个积极但微弱的预测指标,其有用性取决于具体情况。自我一致性是正确性的条件代理,而非独立的置信分数。我们还公开了去识别的每行数据和答案分布。
英文摘要
LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study: 53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME -- 265,000 samples. Using majority-correctness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (rho 0.20-0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst -- over-confident yet no more accurate -- for the most consistent frontier model (agreement >=0.8 on 77% of GPQA case-result entries, 48% of those wrong). An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginal-preserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the de-identified per-run rows and answer distributions.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。