arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型推理的测试时校准学习

Test-time Calibration Learning for Large Language Model Reasoning

Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han

arXiv 2610.02695首次发表:更新:

发表机构

Hong Kong Baptist University(香港浸会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出测试时校准学习(TTCL),一种无标签框架,在未标注目标数据上联合优化推理准确性和置信度校准,显著提升模型性能。

AI 中文摘要

可靠的大语言模型(LLM)不仅需要产生准确的答案,还需要表达能够忠实反映其正确概率的置信度。这种校准对于识别不确定的预测以及支持实际部署中的可靠决策至关重要。近期研究将校准学习纳入强化学习(RL),利用真实正确性监督联合优化答案正确性和口头化置信度。然而,它们对标注数据的依赖限制了其在实际测试时场景中的适用性,在这些场景中,真实标签不可用,且校准可能需要适应新遇到的目标任务。为解决这一挑战,我们提出了测试时校准学习(TTCL),一个无标签框架,直接在未标注的目标任务数据上联合调整推理准确性和口头化置信度。具体而言,TTCL从多个模型生成的响应中为正确性和校准导出自我监督信号,从而在没有真实标签的情况下实现测试时的校准学习。理论分析进一步将TTCL确立为理想校准目标的有界替代。在数学推理和事实问答上的大量实验表明,TTCL在不同模型和任务上持续提高了准确性和校准。在基础模型上,TTCL在八个基准上实现了平均相对准确性提升+40.13%和ECE降低+70.80%。此外,在领域偏移下,TTCL可以进一步提高已经校准模型的准确性和校准,特别是当源域校准对目标任务迁移效果不佳时。在数学到事实问答(math-to-factQA)设置中,TTCL实现了平均相对准确性增益+20.35%并将ECE降低+53.83%。源代码已在此https URL发布。

英文摘要

Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time Calibration Learning (TTCL), a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data. Specifically, TTCL derives self-supervision signals for both correctness and calibration from multiple model-generated responses, enabling calibration learning at test time without ground-truth labels. Theoretical analysis further establishes TTCL as a bounded surrogate for the ideal calibration objective. Extensive experiments on mathematical reasoning and factual question answering demonstrate that TTCL consistently improves both accuracy and calibration across diverse models and tasks. On base models, TTCL achieves an average relative accuracy improvement of +40.13% and an ECE reduction of +70.80% across eight benchmarks. Moreover, TTCL can further improve both accuracy and calibration for already calibrated models under domain shift, particularly when source-domain calibration transfers poorly to target tasks. In the math-to-factQA setting, TTCL achieves an average relative accuracy gain of +20.35% and reduces ECE by +53.83%. The source code is released at https://github.com/tmlr-group/TTCL.

Comments34 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑