发表机构
Microsoft Research Asia(微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨了四个前沿大模型在15个基准上的元认知监控局限性,发现高解题准确率可与弱错误区分能力共存,且自我复核和同伴监督难以改善困难问题及共有错误的置信度评估。
AI 中文摘要
可靠的决策依赖于识别答案可能出错的能力。在生物认知中,元认知监控可以与任务表现分离,这引发了一个问题:在语言模型中,解题与判断之间的联系有多紧密。在此,我们研究了四个前沿模型在15个基准上的置信度报告。高任务准确率可以与较弱的错误区分能力共存:一个模型解决了97%的竞赛数学问题,而其答题时的置信度将正确答案排在错误之上仅勉强优于随机水平。置信度在由独立参考模型解决的问题上更能有效区分正确答案与错误,而复核在参考模型难以解决的问题上带来的改进有限。聚合区分度也奖励将简单问题上的正确答案排在困难问题上的错误之上,而仅基于问题的预测已经能很好地做到这一点。交叉评估在评估者正确回答的情况下帮助最大,而两个模型共有的错误通常仍保持高置信度。困难问题和共有错误仍然是提示式自我复核和同伴监督的难以攻克的目标,即使在具有强大解题能力的模型中也是如此。
英文摘要
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.