发表机构
Johannes Gutenberg University Mainz; University of Colorado Boulder(美因茨约翰内斯·古腾堡大学; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文指出LLM评估中校准指标采纳不足的问题,主张将校准作为一等标准,与性能指标配对报告,以提升模型可信度。
AI 中文摘要
语言模型的校准——即表达或隐含的置信度与经验正确性之间的一致性——是NLP中一个研究充分的子领域。测量校准的方法已经存在。问题在于采纳:在该子领域之外,NLP研究经常引入新模型、数据集和基准,却不检查模型的置信度分数是否有意义。我们认为,这种采纳差距是可信LLM评估的主要障碍。校准不良在两个不同领域造成问题:在部署时,过度自信的错误会造成实际危害;在研究流程内部,诸如LLM-as-a-judge、合成数据生成和主动学习等方法依赖校准置信度却未加以验证。标准校准指标每个示例仅需两个输入:置信度分数和正确性判断。当今使用的大多数基准已经提供这两者,意味着校准可以立即报告。然而,对于开放式生成,定义这两个输入仍是一个开放挑战。我们主张每个NLP子领域应将其主要性能指标与校准分数配对,并呼吁将校准视为每个模型的基本属性,而非一个冷门话题。
英文摘要
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
CommentsAccepted to the 3rd Workshop on Uncertainty-Aware NLP (UncertaiNLP) at EMNLP 2026