arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有释义感知评分的共形预测用于大语言模型不确定性量化

Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

Jiayi Xin, Evan Qiang, Zihan Zhu, Xiang Li, Weijie J. Su, Qi Long

arXiv 2610.04239首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型在语义等价改写下不确定性量化不稳定的问题,提出释义感知评分框架,通过代理模型聚合释义预测构建非一致性评分,实现稳健的共形预测,在多个基准上保持覆盖并减小集合大小。

AI 中文摘要

大语言模型(LLM)的不确定性量化(UQ)旨在提供可靠的预测置信度度量,然而当前方法在保持语义的扰动下往往不稳定。语义等价的释义可能引起预测置信度的显著变化,即使对于具有形式保证的方法(如共形预测)也是如此。为解决此问题,我们提出一种对语义改写具有鲁棒性的释义感知UQ框架。我们的方法在LLM隐藏状态上训练一个轻量级代理模型,并跨释义聚合其预测以构建标签级非一致性评分。在分数可交换性下,共形校准保持边际覆盖。该保证在仅测试改写下也能成立,前提是释义流程满足额外的分布对齐条件。我们评估了三种设置(正常、完全改写和半改写),分别对两个数据集均不应用改写、对校准和测试数据集均应用改写、或仅对测试数据集应用改写。在七个多项选择问答基准和多个模型家族上,我们的方法产生紧凑的预测集,经验覆盖通常接近名义目标,即使在半改写设置中也是如此。消融研究表明,学习的代理模型贡献了集合大小减少的大部分,而释义增强训练和推理时聚合提高了改写下的稳定性。代码可在以下URL获取。

英文摘要

Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings. Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores. Under score exchangeability, conformal calibration retains marginal coverage. This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition. We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively. Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting. Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording. Code is available at https://github.com/Raina-Xin/PA_Score.

CommentsNeurIPS 2026 Poster

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑