发表机构
University of Stuttgart; Robert Bosch GmbH; University of Oxford; LMU Munich; Tsinghua University; University of Southampton; University of Oslo(斯图加特大学; 罗伯特·博世有限公司; 牛津大学; 慕尼黑大学; 清华大学; 南安普顿大学; 奥斯陆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对配置偏移对共形大语言模型有效性的影响展开系统分析,通过多维度实证研究揭示其削弱覆盖率的问题,提出两种实用缓解措施以恢复覆盖率并保留效率。
AI 中文摘要
共形预测(Conformal Prediction, CP)是一种用于不确定性量化的无分布框架,近期已被适配到大语言模型(Large Language Models, LLMs)中,可在可交换性下提供具有有限样本覆盖保证的预测集。然而对于LLMs而言,非一致性分数通常由推理流程而非仅固定模型产生,因此它们不仅依赖数据分布,还受提示模板、解码参数、部署设置等可配置因素影响。由于这些配置在实际应用中会被频繁修改,但很少被视为一种偏移来源,它们对CP有效性的影响仍未被充分理解。我们将此称为“配置偏移”,并沿提示模板、解码温度、权重量化三个维度对其进行系统研究。在涵盖9个LLMs、4个数据集和4个非一致性分数的广泛实证研究中,我们发现配置偏移会持续削弱CP有效性,常使经验覆盖率低于目标值。相比之下,效率在很大程度上得以保留:有效预测集的大小与独立同分布(i.i.d.)基线接近。我们推导了覆盖下界,将这种损失归因于校准与测试分数分布之间的差异,并将其有限样本插件版本作为偏移严重程度的经验诊断指标。我们进一步表明,这些发现可用于实际缓解措施:受边界启发的重新校准在测试样本有限时有效,而感知脆弱性的校准集成无需测试数据即可恢复大部分损失的覆盖率。
英文摘要
Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability. Yet for LLMs, nonconformity scores are often induced by an inference pipeline, not just a fixed model, making them depend not only on the data distribution but also on configurable factors such as the prompt template, decoding parameters, and deployment setting. Since such configurations are routinely modified in practice but rarely treated as a source of shift, their impact on CP validity remains poorly understood. We call this \emph{configuration shift} and study it systematically along three axes: prompt template, decoding temperature, and weight quantization. In a broad empirical study spanning $9$ LLMs, $4$ datasets, and $4$ nonconformity scores, we find that configuration shift consistently erodes CP validity, often driving empirical coverage below the target. By contrast, efficiency is largely preserved: valid prediction sets remain close in size to the i.i.d. baseline. We derive coverage lower bounds that attribute this loss to a discrepancy between calibration and test score distributions, and use their finite-sample plug-in versions as empirical diagnostics of shift severity. We further show that these findings lead to practical mitigations: bound-inspired recalibration is effective with limited test examples, while fragility-aware calibration ensembling recovers much of the lost coverage without test data.
CommentsAccepted to EMNLP 2026 as a main conference paper