评估在不断演化的模型知识下的持久校准
Evaluating Persistent Calibration under Evolving Model Knowledge
浏览论文内容
中文总结 AI 辅助
本研究提出持久校准问题,评估置信度估计器在模型知识变化时的泛化能力,发现现有方法在对比集上校准不足,多检查点训练可改善校准。
中文摘要 AI 辅助
随着人工智能系统从静态存储库转变为能够持续适应和学习的智能体,保持其可信度意味着需要为支撑它们的模型配备能够动态反映其不断变化的技能和知识的置信度估计能力。我们引入了持久校准问题,该问题要求置信度估计器能够忠实地反映模型中所包含的知识,随着知识的变化而变化,且无需重复监督。我们通过检查开放模型不同检查点上的持久校准来操作化这一问题,询问在较早检查点上训练的置信度估计器能否泛化到较晚的检查点。具体而言,我们旨在阐明置信度是否依赖于知识,这一问题对置信度估计的可靠性具有影响。为了衡量这种关系,我们在知识对比集上定义并评估校准:这些子集包含一个检查点正确回答而另一个检查点错误回答的问题,反映了知识的变化。我们表明,与在未来检查点上训练的神谕方法相比,推理时和微调方法在对比集校准上均表现不佳,即使对于在全数据集上校准良好的方法也是如此。我们为持久校准具有挑战性这一假设提供了证据,因为在给定检查点上存在大量校准良好的可能置信度函数,其中只有一些依赖于能够泛化到其他检查点的元知识特征。为了改进对比集校准,我们表明多检查点训练有所帮助,这为识别在变化知识中保持稳健的置信度特征提供了一条途径。
英文摘要
As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.
发表机构
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。