发表机构
University of Science and Technology of China; Qwen Business Unit of Alibaba; National University of Singapore(中国科学技术大学; 阿里巴巴通义千问事业部; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RLVR模型过度自信问题,提出CREDO方法,用可微读出替代文本采样置信度,并通过回归训练及加权轨迹,在数学和代码推理中同时提升准确性与校准度。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)训练推理模型产生正确答案,但并不确保其陈述的置信度是经过校准的。由此产生的模型系统性地过度自信。近期方法通过在RLVR循环内让模型在答案旁陈述一个数值置信度来训练校准,但它们都通过将置信度作为文本采样来获得。这一选择带来两个代价:采样的置信度引入方差,并在实践中坍缩为少数几个不同值;且采样使置信度不可微,迫使校准损失通过标量奖励进行。我们提出CREDO(置信度读出)以确定性读出取代采样。在RLVR优化正确性的同时,CREDO从模型输出分布中的专用token对读取置信度,并通过可微回归进行训练。CREDO进一步将训练后的置信度转化为准确性的信号,根据置信度与结果不一致的程度对轨迹进行加权,从而使准确性和校准共同提升。在数学和代码推理中,CREDO达到了最佳的准确性和校准效果,且这些增益扩展到弃权(不执行)和选择性预测。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.