arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33290cs.CL

校准,而非答案选择:在推理模型中蒸馏内部置信度

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

针对推理模型口头置信度过度自信的问题,提出探针引导的自蒸馏(Probe-SD),利用隐藏状态线性探针校准置信度并微调模型,显著降低ECE,优于事后重校准等方法。

中文摘要 AI 辅助

使用二元正确性奖励的强化学习训练的是正确性,而非校准的置信度。推理模型口头表达的置信度系统性地过度自信,且问题不仅仅是规模问题:口头表达的置信度反映的是模型愿意对某个答案做出承诺的程度,而非该答案正确的可能性。因此,事后重新缩放只能拟合一种分布,却很少能迁移。我们转而审视模型内部。在事实性问答任务上,对思维链与答案之间的隐藏状态进行线性探针,其校准效果显著更好:在四个基准和两个模型家族上,其期望校准误差比口头表达分数低5到38倍。然而,当用于从N个采样答案中进行选择时,同一探针几乎与多数投票持平,却远不及预言机。内部状态能很好地回答“我有多确定”,却很难回答“哪个答案是正确的”,因此该信号应作为置信度来报告,而非用于选择答案。由此,我们提出了探针引导的自蒸馏(Probe-SD):用探针对模型自身采样的轨迹进行评分,覆盖每条轨迹所陈述的置信度,并对同一家族的基础检查点进行微调,从而在测试时仅保留模型本身。在Qwen3-14B上,Probe-SD将域内ECE从0.178降至0.024,将域外ECE从0.542降至0.113,同时其表现也优于事后重新校准和自一致性蒸馏。由此得到的置信度校准良好,且可用于加权投票,这些行为此前被归因于在线强化学习,而这里仅通过监督微调即可获得。

英文摘要

Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.

↑