arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33886cs.CL

LLMs在训练预测自身准确性时学习不同形式的元认知

LLMs learn different forms of metacognition when trained to predict their own accuracy

发表机构LNC2,法国国家健康与医学研究院 · DEC,巴黎高等师范学院,巴黎文理研究大学 · Flowers AI与认知科学实验室
另 1 家 · 查看机构详情
  • LNC2, INSERM(LNC2,法国国家健康与医学研究院)
  • DEC, ENS, PSL(DEC,巴黎高等师范学院,巴黎文理研究大学)
  • Flowers AI & CogSci Lab(Flowers AI与认知科学实验室)
  • Centre Inria de l’Université de Bordeaux(波尔多大学Inria中心)

机构由 AI 辅助整理,请以论文原文为准。

Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过训练10个开放权重LLM预测自身准确性,发现其置信度反映两种信号:接近训练数据时追踪真实准确性,其他领域则追踪输出一致性,表明校准训练可能无法教会模型普遍检测自信错误。

中文摘要 AI 辅助

大型语言模型被训练为无论是否拥有相关知识都总是产生答案,这导致它们捏造事实。先前的研究表明,大型语言模型的置信度估计与其实际表现对应不佳,而微调可以大幅改善这种对应关系。然而,模型在此类训练中实际学到了什么仍鲜为人知。我们通过训练10个开放权重的大型语言模型在回答事实性多项选择题之前预测自身准确性,来研究它们如何获得元认知监控能力,即知道自己知道什么的能力。我们发现,训练后的置信度反映了两种不同的信号。在接近训练数据的问题上,它追踪模型的真实准确性;而在其他领域,它转而追踪输出一致性:模型答案分布的集中程度。输出一致性追踪在训练早期出现,并跨数据集泛化,而准确性追踪则发展较晚,且仍局限于训练分布。这些结果表明,校准训练可能不会教会模型普遍检测它们自信犯下的错误,并引发了关于人工系统中元认知本质的更广泛问题。

英文摘要

Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.

补充信息

↑