大语言模型中的稳定校准误差:高置信度错误的实用视角
Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
浏览论文内容
中文总结 AI 辅助
本研究提出稳定校准误差概念,结合两类诊断方法,发现自我批判提示可降低模型隐藏状态敏感性,但高置信度错误未必是推理脆弱,部分或为稳定校准不当。
中文摘要 AI 辅助
大语言模型的高置信度错误常被视为内部推理脆弱性的证据。本研究探讨了另一种可能性:稳定校准误差,即错误答案在小扰动下仍保持局部稳定。我们结合两类诊断方法:一是标签感知的输出级审计分数,该分数在强制回答基准下,按置信度变化和过度置信错误对领域进行排序;二是内部敏感性探测,用于测量隐藏状态的变化。在多领域二元事实审计集上,该审计分数可追踪感知弃权(不执行)的自我批判减少决策损失的情况,尽管直接标记基准对该收益的排序更强。内部层面,自我批判提示在三个开放权重模型的各层中均持续降低隐藏状态敏感性,这支持提示诱导的局部稳定而非纯输出级弃权模式,但不意味着校准得到改善:审计定义的过度置信错误并不比置信正确答案明显更具局部敏感性,因此部分高置信度错误可能是稳定且校准不当,而非单纯脆弱。
英文摘要
High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly. Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.
发表机构
- ToppyMicroServices OÜ
机构由 AI 辅助整理,请以论文原文为准。