发表机构
State Key Laboratory of AI Safety; Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences(人工智能安全国家重点实验室; 中国科学院计算技术研究所; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型错误答案仍高置信的问题,提出CoCal方法,通过共享轨迹但分离参数与优化机制训练轻量伴生模型进行置信度校准,在不损失任务性能下提升估计可靠性,并在多规模实验中验证泛化性。
AI 中文摘要
可靠的自我评估对于大语言模型(LLMs)至关重要,然而当它们的答案错误时,它们往往仍然保持高度自信。我们研究“并发置信度校准”,即在能力提升的同时学习置信度,而不是仅在训练后进行校准。基于可验证奖励的强化学习(RLVR)为该范式提供了自然的设置,因为它持续产生带有可验证正确性反馈的响应。然而,现有的并发方法通过共享策略参数内的强化学习同时学习能力和置信度,可能将两个根本不同的学习问题耦合在一起。我们反而提出“共享经验但分别学习”:能力和置信度从相同的轨迹中学习,但通过独立的优化机制和参数进行。基于这一原则,我们引入了CoCal(伴生置信度校准),它从 rollout 隐藏状态和验证器派生的正确性监督中训练一个轻量级伴生模型,同时保持任务优化不变。在Qwen3-8B和Qwen3-14B上的实验表明,CoCal在不牺牲任务性能的情况下改善了置信度估计,优于基于强化学习的并发方法和匹配的事后校准。学习到的伴生模型进一步跨领域和策略变化进行泛化,而CoCal的益处在这两种规模下均持续存在。
英文摘要
Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.