论言语化置信度先验在大型推理模型校准中的陷阱
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
AI总结:
本文揭示大型推理模型置信度先验高度集中导致校准困难,提出CalibSFT监督微调阶段以塑造广泛支持的置信度先验,在16个基准上提升校准性能并保持准确性。
AI中文摘要:
大型推理模型(LRMs)在表达其不确定性时常常过于自信。置信度感知的强化学习(RL)为优化校准提供了一种有前景的方法。然而,它依赖于策略内(on-policy)采样,因此受限于模型在RL之前的置信度分布,我们称之为置信度先验。在这项工作中,我们揭示了现成的LRM表现出一种置信度先验,该先验高度集中于少数几个高值,并且这种集中在整个RL过程中持续存在。理论上,我们证明了这种集中抑制了针对罕见采样置信度值的策略梯度更新,并抬高了期望Brier风险的下界。为了克服这一探索瓶颈,我们提出了CalibSFT,一个即插即用的监督微调阶段,在RL之前塑造一个具有广泛支持的校准置信度先验。对于每个问题,CalibSFT构建置信度目标,将其成功率与响应级别的正确性相结合,这被证明能保持适当评分的优化性,然后跨置信度谱平衡训练响应,以在RL期间实现多样化的置信度探索。为了从不正确的响应中学习而不模仿其推理,CalibSFT引入了正确性条件监督,指导所有响应的置信度,同时仅在正确的响应上监督推理。在16个数学和通用推理基准上,将CalibSFT纳入五种代表性RL算法中,减少了校准误差并提高了判别能力,同时保持了相当的准确性。此外,CalibSFT为下游选择性预测和模型路由带来了实际益处。我们的代码可在以下网址获取:https://this URL。
英文摘要:
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term confidence prior. In this work, we reveal that off-the-shelf LRMs exhibit a confidence prior heavily concentrated on a few high values, which persists throughout RL. Theoretically, we prove that this concentration suppresses policy gradient updates for rarely sampled confidence values and inflates the lower bound on expected Brier risk. To overcome this exploration bottleneck, we propose CalibSFT, a plug-and-play supervised fine-tuning stage that shapes a calibrated confidence prior with broad support before RL. For each question, CalibSFT constructs confidence targets combining its success rate with response-level correctness, which provably preserves proper-scoring optimality, and then balances training responses across the confidence spectrum to enable diverse confidence exploration during RL. To learn from incorrect responses without imitating their reasoning, CalibSFT introduces correctness-conditional supervision, guiding confidence across all responses while supervising reasoning only on correct ones. Across 16 mathematical and general reasoning benchmarks, incorporating CalibSFT reduces calibration errors and improves discrimination across five representative RL algorithms while preserving comparable accuracy. Furthermore, CalibSFT delivers practical benefits for downstream selective prediction and model routing. Our code is available at https://github.com/ml-stat-Sustech/verbalized-confidence-training.