AI 中文总结
该研究在隐藏信息国际象棋变体中发现,大语言模型的高置信度行动存在严重信念校准偏差,传统评估指标与信念质量脱节,仅结果评估无法检测此问题。
AI 中文摘要
智能体系统越来越多地基于模型自身声明的置信度来决定行动,这一机制假设置信度在行动时刻能反映正确性。我们在一种隐藏信息的国际象棋变体中验证了这一点:王/后身份可被秘密、反复地在棋子间转移,且智能体每回合都会单独声明其关于对手隐藏王/后棋子的概率分布,该分布将与游戏结束后可恢复的真实情况进行比对评分。在两个独立批次中,基于对隐藏棋子位置的高声明置信度(≥0.5)实施的吃子操作,62次中仅1次正确。校准偏差几乎完全集中在这些事件中:原始批次中99.3%的偏差源于此,复现批次中为98.7%。该模式以较弱形式在四个额外模型配置中持续存在,这些配置涵盖了第二个模型提供商——此为该发现的适用范围,而非能力预测校准的证据:在固定外部排行榜分数下的同模型比较显示,仅思考预算的变化就使该指标的变动幅度几乎与跨模型的巨大差距相当。在另一维度,传统评估指标(合法性、成本、延迟、完成率)可与信念质量完全脱节,在所有传统指标上表现最佳的配置,其信念质量却是所测试配置中最差的。呈现此模式的模型仍可赢得其信念所针对的游戏,这正是仅基于结果的评估无法检测到该问题的原因。
英文摘要
Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\geq 0.5$) about the hidden piece's location were correct in 1 of 62 cases. The calibration deficit is concentrated almost entirely in these events: 99.3% of it in the original batch, 98.7% in the replication. The same pattern, in weaker form, orders consistently (point estimates only; most pairwise gaps are not statistically distinguishable at this sample size) across four further model configurations spanning a second provider -- reported as scope for the finding, not as evidence that capability predicts calibration: a same-model comparison at a fixed external leaderboard score shows a deliberation-budget change alone moves the metric by nearly as much as a large cross-model gap. In a separate seat, conventional evaluation axes -- legality, cost, latency, completion rate -- can dissociate entirely from belief quality, with the configuration winning on every conventional axis producing the worst belief quality tested. A model exhibiting this pattern can still win the game its belief was about, which is why outcome-only evaluation would not detect it.