发表机构
DreamAI(DreamAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析语言模型免训练置信度信号,发现直接口头表达优于多样本一致性,且自报告对提示方式、相关错误和基准噪声敏感。
AI 中文摘要
大型语言模型可以在生成内容的同时报告一个数值置信度,但尚不清楚这种报告是否仅仅是经过校准的修辞。我们分析了两种模型家族在相同100个TriviaQA问题上的三种免训练信号:与答案一起口头表达的置信度、事后$P(\mathrm{True})$以及三次额外生成的一致性。直接口头表达是一个出人意料的强基线:在审计基准错误后,其正确性预测的AUROC达到0.956和0.937。三样本一致性明显较弱(0.765和0.790),而固定插值与口头置信度的组合没有统计上可靠的收益。一个模型的九个错误中有四个,另一个模型的八个错误中有两个获得了样本的一致支持,表明自一致性可能放大共享的误解。使用等效提示对相同固定答案重新获取置信度,平均分数变化为0.043至0.084,并在0.8阈值下翻转了4%至9%的决策。对100个带置信度标签的传记声明的探索性审计进一步发现,支持与矛盾声明之间的置信度差距仅适中。这些结果表明,有用的自我报告仍然对提示方式、相关错误和基准噪声敏感。
英文摘要
Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4\% to 9\% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.
Commentsworkshop