arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

确定性不仅仅是正确性:重新思考LLM推理中的词元级确定性

Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li

arXiv 2610.00296首次发表:更新:

发表机构

Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究重新评估LLM推理中词元级确定性的预测能力,发现其在不同预测目标上表现不一,并据此提出一种利用确定性分配响应数量和加权投票的方法,在提升准确率的同时大幅降低生成成本。

AI 中文摘要

词元级确定性在LLM训练和推理中被广泛用作正确性的代理指标。然而,基于确定性的方法的性能既取决于确定性分数中的信息,也取决于这些分数的使用方式。因此,我们通过跨模型和任务的受控实证评估,直接评估确定性的预测能力。我们区分两个预测目标:识别模型更可能正确回答的问题,以及区分同一问题的正确与错误回答。在我们的实验中,确定性在识别模型可能正确回答的问题方面通常优于区分同一问题的正确与错误回答。确定性还随词元类型和词内位置系统性地变化,反映了词和文本形式的局部属性。关于问题难度的信息在生成早期出现,而关于回答正确性的较弱信息则更集中在接近末尾处。这些发现表明,确定性为决策提供的信息取决于预测目标、模型、确定性度量以及响应中哪些词元位置被纳入聚合。我们进一步展示了这些发现在测试时计算中的实用价值。我们使用生成早期的确定性来分配响应数量,并使用每个响应接近末尾的确定性来加权答案投票。与固定采样多数投票基线相比,该方法将整体准确率从78.71%提升至79.54%,同时将生成词元成本降低了82.4%。

英文摘要

Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answer correctly and distinguishing correct from incorrect responses to the same question. In our experiments, certainty is generally better at identifying questions a model is likely to answer correctly than at distinguishing correct from incorrect responses to the same question. Certainty also varies systematically across token types and positions within words, reflecting local properties of words and text form. Information about question difficulty appears early in generation, while the weaker information about answer correctness is more concentrated near the end. These findings show that the information certainty provides for decisions depends on the prediction target, the model, the certainty metric, and which token positions in the response are included in aggregation. We further demonstrate the practical value of these findings for test-time compute. We allocate the number of responses using certainty early in generation and weight answer votes using certainty near the end of each response. Compared with a fixed-sampling majority-voting baseline, this approach increases overall accuracy from 78.71\% to 79.54\% while reducing generated-token cost by 82.4\%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑