发表机构
University of Exeter; Cardiff University; University of the West of England; University of Macau(埃克塞特大学; 卡迪夫大学; 西英格兰大学; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM作为评估者时候选响应增多导致置信度假设失效的问题,提出“先定位后决策”框架,结合共形预测与校准规则,提升保证成功率与覆盖率。
AI 中文摘要
大型语言模型(LLM)越来越多地被用作评估者,以评估输出质量和偏好对齐,但提供与人类判断一致的可靠保证仍然具有挑战性。最近的研究引入了置信度阈值方法,这些方法为成对比较提供此类保证,依赖于“估计置信度越高,与人类的分歧风险越低”的假设。然而,当候选响应数量增加时,这一假设可能不成立,因为在众多备选方案中分配概率质量会扭曲置信度估计。为解决该问题,我们提出“先定位后决策”框架:首先,采用共形预测定位一个小型候选列表,该列表以高概率包含人类偏好的响应;然后,使用校准的基于置信度的规则从该列表中选择性选择单个响应或弃权(不执行)。该设计恢复了置信度与分歧风险之间的单调关系,并实现了高概率的一致保证。在多个数据集和多个候选大小上,使用裁判LLM进行的实验表明,我们的框架始终比单阶段基线实现更高的保证成功率和显著更高的覆盖率。
英文摘要
Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimates. To address this issue, we propose a Localize-Then-Decide framework. First, conformal prediction localizes a small shortlist that contains the human-preferred response with high probability. Then, a calibrated confidence-based rule selectively chooses a single response from this shortlist or abstains. This design restores the monotonic relationship between confidence and disagreement risk and enables high-probability agreement guarantees. Experiments with multiple candidate sizes across several datasets and judge LLMs demonstrate that our framework consistently achieves higher guarantee success rates and substantially higher coverage than single-stage baselines.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Code: https://github.com/llm2409/Localize-Then-Decide