重新思考LLM-as-a-Judge中的言语化置信度:2025年后专有模型上的兼容性转变
Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models
- Cisco Systems(思科系统)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发现2025年后专有模型上言语化置信度优于对数概率,提出过度自信建议和自辩论两个新要素,改善校准与主观性稳健性,建议LLM-as-a-Judge广泛采用软评分。
AI中文摘要:
言语化置信度长期以来被视为过度自信、粗糙且倾向于整数聚集,如今在顶级专有模型上已成为LLM-as-a-Judge更稳健的软评分机制。在SummEval、AggreFact和HelpSteer2数据集上,涵盖多达18个LLM,我们表明在2025年后的模型上,偏好对数概率的标准建议不再成立,而言语化置信度是更好的信号。我们称之为兼容性转变。在标准言语化置信度基线之上,我们引入了两个新要素:过度自信建议和自辩论。它们共同改善了校准、评分分布扩展以及对任务主观性的稳健性。我们进一步观察到一种生成效应:2025年后的模型在平衡准确率损失很小的情况下适应了这两个新增要素,而2025年前的模型则付出了可衡量的代价。与基于对数概率的G-Eval相比,言语化置信度在GPT系列顶级版本上是更具主观性稳健性的软信号。这种转变在仅报告准确率的情况下是不可见的。我们建议在LLM-as-a-Judge中更广泛地使用软评分,而不是默认使用硬预测。更广泛地说,言语化置信度已从对数概率的较弱替代品转变为当代LLM评判者的实用软评分机制。
英文摘要:
Verbalized confidence, long dismissed as overconfident, coarse, and prone to round-number clustering, is now the more robust soft-scoring mechanism for LLM-as-a-Judge on top-tier proprietary models. Across SummEval, AggreFact, and HelpSteer2, spanning up to 18 LLMs, we show that the standard advice to prefer log-probabilities no longer holds on post-2025 models, where verbalized confidence is the better signal. We call this a compatibility shift. On top of a standard verbalized-confidence baseline, we introduce two new ingredients: an overconfidence advisory and self-debate. Together they improve calibration, score-distribution spread, and robustness to task subjectivity. We further observe a generation effect: post-2025 models accommodate these two additions with little balanced-accuracy cost, whereas pre-2025 models pay a measurable penalty. Compared with logprob-based G-Eval, verbalized confidence is the more subjectivity-robust soft signal on GPT-family top-tier releases. The shift is invisible under accuracy-only reporting. Rather than defaulting to hard predictions, we recommend broader use of soft scoring in LLM-as-a-Judge. More broadly, verbalized confidence has moved from a weaker substitute for logprobs to a practical soft-scoring mechanism for contemporary LLM judges.