发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TRACE通过单遍解码轨迹记录令牌级不确定性并应用局部风险算子,将局部风险转化为答案级置信度,显著提升校准性能。
AI 中文摘要
可靠的置信度估计对于大型语言模型的部署至关重要。然而,答案级别的校准仍然具有挑战性,因为生成错误往往是局部性的:一个响应可能在整体上流畅且具有高概率,但在关键数字、实体或事实性声明上仍然失败。现有的估计器将令牌概率、序列似然、熵或波束统计量压缩为全局分数,这可能会稀释此类局部风险信号。我们提出TRACE,一种单遍、保留解码答案的置信度估计器,将解码时的不确定性视为一个三步轨迹:(i)在解码过程中记录令牌级别的惊讶度和预测熵,(ii)应用局部风险算子以保留不确定性尖峰,(iii)将局部化的轨迹风险转换为答案级别的置信度。TRACE产生无标签的风险分数,而TRACE+使用保留分割将仅轨迹特征校准为概率,无需额外生成或外部验证器。我们针对19个校准基线评估了四项任务,TRACE+将Brier分数从0.149降至0.137,并将AUROC从0.758提升至0.792,优于最强似然基线。在七个LLM中,TRACE+将最佳非TRACE基线池的Brier分数从0.136改善至0.120,AUROC从0.764提升至0.817。结果表明,局部化解码时风险为校准提供了一种通用方法。
英文摘要
Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a critical number, entity, or factual claim. Existing estimators compress token probabilities, sequence likelihoods, entropy, or beam statistics into a global score, which can dilute such local risk signals. We propose TRACE, a single-pass, decoded-answer-preserving confidence estimator that treats decoding-time uncertainty as a trajectory through three steps: (i) recording token-level surprisal and predictive entropy during decoding, (ii) applying local risk operators to preserve uncertainty spikes, and (iii) converting localized trace risk into answer-level confidence. TRACE produces a label-free risk score, while TRACE+ calibrates trace-only features into probabilities using a held-out split, without extra generations or external verifiers. We evaluate four tasks against 19 calibration baselines, and TRACE+ reduces Brier from 0.149 to 0.137 and improves AUROC from 0.758 to 0.792 over the strongest likelihood baseline. Across seven LLMs, TRACE+ improves over the best non-TRACE baseline pool from 0.136 to 0.120 Brier and from 0.764 to 0.817 AUROC. Results show that localizing decoding-time risk provides a general approach to calibration.
CommentsEMNLP 2026 Findings