arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24303cs.LGcs.CL

SupportCal:通过参考支持与佐证对训练后大语言模型进行无标签校准

SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration

Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, José Miguel Hernández-Lobato, Junbin Gao

首次发表
浏览论文内容

中文总结 AI 辅助

针对训练后模型过度自信问题,提出无标签校准方法SupportCal,利用预训练参考模型的支持与佐证对不一致样本加权,在多个基准上降低期望校准误差。

中文摘要 AI 辅助

训练后阶段通常能提升任务性能,但可能损害置信度校准,导致训练后语言模型(PoLMs)比其对应的预训练语言模型(PLMs)更加过度自信。由于任务特定的有标签校准数据可能成本高昂或不可用,对应的预训练PLM为事后校准提供了天然的无标签参考。先前的基于一致性门控的PLM参考校准方法仅使用PoLM与其PLM参考一致的样本来拟合标量温度,排除不一致样本,因为直接对齐可能使拟合温度过高并导致欠自信。我们重新审视这种二元处理方式。一项受控的重新引入诊断揭示了非单调的总体效应:纳入中等比例的不一致样本可以改善校准,而当单位权重纳入接近完整不一致集时,益处逐渐减弱。我们提出SupportCal,一种无标签的事后方法,该方法以单位权重保留一致样本,并根据从规模兼容的候选池中选取的预训练参考的自身基础PLM相对支持与佐证,为不一致样本分配连续权重。我们进一步刻画了所得加权目标何时允许有限最优温度。在MedMCQA和MathQA上,对于几乎所有评估的目标模型配置,SupportCal比仅一致基线实现了更低的ECE;补充的TweetEval情感分析结果在固定标签分类任务上显示出相同模式。

英文摘要

Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower mean ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.

发表机构

  • The University of Sydney(悉尼大学)
  • University of Cambridge(剑桥大学)
  • The University of Adelaide(阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑