AI 中文总结
本研究提出一种结合结构化推理与感知距离校准技术的校准反思方法,通过三项创新提升LLM置信度估计,在多数据集上验证了其在两类任务中的有效性,助力LLM可靠置信度估计的发展。
AI 中文摘要
部署大语言模型(LLM)面临的一个关键挑战是开发可靠的置信度估计机制,使系统能够判断何时信任模型输出、何时寻求人工干预。本文提出一种用于增强LLM置信度估计的校准反思方法,该框架结合了结构化推理与感知距离的校准技术。该方法引入三项关键创新:(1)最大置信度选择(MCS)方法,可全面评估所有可能标签的置信度;(2)基于反思的提示机制,可提升推理可靠性;(3)感知距离的校准技术,可考虑标签间的序数关系。我们在HelpSteer2、Llama T-REx及一个私有对话数据集等多样化数据集上对该框架进行评估,证明其在对话类和基于事实的分类任务中均有效。本研究为开发可靠且校准良好的LLM置信度估计方法这一更广泛目标作出贡献,支持关于模型信任与人工判断的明智决策。
英文摘要
A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate their confidence, enabling systems to determine when to trust model outputs versus seek human intervention. We present a Calibrated Reflection approach for enhancing confidence estimation in LLMs, a framework that combines structured reasoning with distance-aware calibration technique. Our approach introduces three key innovations: (1) a Maximum Confidence Selection (MCS) method that comprehensively evaluates confidence across all possible labels, (2) a reflection-based prompting mechanism that enhances reasoning reliability, and (3) a distance-aware calibration technique that accounts for ordinal relationships between labels. We evaluate our framework on diverse datasets, including HelpSteer2, Llama T-REx, and a proprietary conversational dataset, demonstrating its effectiveness across both conversational and fact-based classification tasks. This work contributes to the broader goal of developing reliable and well-calibrated confidence estimation methods for LLMs, enabling informed decisions about model trust and human judgement.
CommentsPublished at TrustNLP 2025 (NAACL 2025 Workshop)
Journal refProceedings of the 5th Workshop on Trustworthy Natural Language Processing (TrustNLP 2025), pages 399-411, Albuquerque, New Mexico. Association for Computational Linguistics, 2025
DOI:10.18653/v1/2025.trustnlp-main.26