发表机构
Transformer Lab(Transformer实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究用模型自身免费的置信度信号,替代带标签数据集微调模型使其学会弃权(不执行),发现其效果与带标签监督相当,仅在自信错误事实上存在盲区。
AI 中文摘要
大型语言模型陈述虚假事实时的流畅度与陈述真实事实时相当,但模型内部往往“知道”自己是否处于不稳定状态:对于答错的事实,其分配给自身答案的概率往往会下降。通常,要让模型据此采取行动,即教导模型弃权(不执行)而非猜测,需要包含正确和错误答案的带标签数据集。本文探究模型自身的置信度(免费且无需标签)是否可替代带标签数据集完成该任务。我们对每个模型(采用LoRA)进行微调,仅使用该信号且无需正确性标签,在其冻结置信度高时给出答案,置信度低时说“我不确定”。在短格式事实问答任务中,针对6个开源权重模型(1B至8B,两个系列),由独立评判模型判定正确性,这种无标签方法与带标签监督的弃权微调效果相当:在匹配覆盖率下,两者无统计学上可检测的差异。针对困难示例而非弃权的对照实验无帮助,表明收益来自校准而非死记硬背。该信号的唯一盲区是自信的错误事实,无法被标记。因此,在教导模型何时弃权(不执行)时,模型自身的怀疑是带标签数据集的近乎免费替代品。代码和人工制品可应要求提供。
英文摘要
Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers. We ask whether the model's own confidence, which is free and needs no labels, can do that job instead. We fine-tune each model (with LoRA) to answer when its frozen confidence is high and to say "I'm not sure" when it is low, using the signal alone and no correctness labels. Across six open-weights models (1B-8B, two families) on short-form factual question answering, with correctness adjudicated by an independent judge model, this label-free recipe holds its own against label-supervised abstention-tuning: at matched coverage we find no statistically detectable difference between the two. A control that drills hard examples instead of abstaining does not help, indicating the gain comes from calibration, not rote memorization. The signal's one blind spot is confidently wrong facts, which it cannot flag. A model's own doubt is thus a near-free substitute for a labelled dataset when teaching it when to abstain. Code and artifacts are available on request.