发表机构
Zhangjiang University; Ibaraki University(张江大学; 茨城大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLMs问答的不确定性控制缺陷,提出A-CRC-QA事后校准框架,在两类问答数据集上实现了接受答案可靠性与保留率的良好权衡。
AI 中文摘要
大型语言模型(LLMs)可能生成流畅但错误的答案,因此不确定性量化对于可靠的问答至关重要。然而,启发式不确定性得分无法完美区分正确与错误的预测,直接应用固定不确定性阈值无法对接受答案中的错误率进行统计控制。为解决这一局限,我们提出A-CRC-QA,一种用于不确定性感知选择性问答的事后校准框架。该方法将选择条件下的错误控制重新表述为线性期望约束,并应用受共形风险控制启发的单调经验风险校准程序。由于所得的实例级损失通常关于接受阈值非单调,我们的框架针对渐近而非有限样本的风险控制。A-CRC-QA与模型无关,无需额外训练,可与不同的不确定性估计器结合。在CoQA和MedMCQA上的实验表明,其适用于开放式和封闭式问答,与未校准及基于置信度边界的基线相比,在接受答案的可靠性和答案保留之间实现了良好的权衡。
英文摘要
Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic uncertainty scores cannot perfectly distinguish correct predictions from incorrect ones, and directly applying a fixed uncertainty threshold provides no statistical control over the error rate among accepted answers. To address this limitation, we propose A-CRC-QA, a post-hoc calibration framework for uncertainty-aware selective question answering. The proposed method reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control. Since the resulting instance-wise loss is generally non-monotone with respect to the acceptance threshold, our framework targets asymptotic rather than finite-sample risk control. A-CRC-QA is model-agnostic, requires no additional training, and can be combined with different uncertainty estimators. Experiments on CoQA and MedMCQA demonstrate its applicability to both open-ended and closed-ended question answering, achieving a favorable trade-off between accepted-answer reliability and answer retention compared with uncalibrated and confidence-bound-based baselines.