arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DirEAG:用于校准数学推理中口头表述置信度的Dirichlet证据聚合方法

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

Haorui Xu, Yuzhou Zhu, Liyuan Gao

arXiv 2608.20717首次发表:更新:

发表机构

School of Mathematics, Jilin University; Leicester International Institute, Dalian University of Technology(吉林大学数学学院; 大连理工大学莱斯特国际学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型数学推理中口头表述置信度难以校准的问题,提出DirEAG方法,将置信度观测值转换为校准软证据,在多个数据集和模型上验证其校准效果优于现有方法。

AI 中文摘要

可靠的置信度估计对于将大语言模型应用于数学推理至关重要,但黑盒模型的口头表述置信度难以校准。当同一问题在多个置信度引导提示下被查询时,得到的答案-置信度观测值包含有用的不确定性信息,但这些信息的尺度会因引导水平、模型和数据集而异。现有黑盒不确定性方法通常依赖答案一致性、样本一致性或熵,这些方法描述的是输出变化,但未对自我报告置信度的数值含义进行建模。相反,对提取的置信度直接求平均或进行启发式聚合,无法学习到依赖于提示和任务的偏差。我们提出DirEAG,一种Dirichlet证据聚合方法,该方法将每个提取的答案-置信度观测值转换为对生成的候选答案以及额外空状态的校准软证据,使模型能够表示无候选答案正确的情况。在GSM8K、SVAMP和GSM-Hard数据集上使用Qwen、Mistral和Gemma模型进行的实验表明,与直接置信度平均和启发式置信度引导聚合相比,DirEAG通常能实现更好的校准,同时保持有竞争力的答案选择能力。 ablation研究进一步揭示,证据聚合和最终的二元校准解决了校准问题的不同部分。

英文摘要

Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steering prompts, the resulting answer-confidence observations contain useful uncertainty information, yet their scales may shift across steering levels, models, and datasets. Existing black-box uncertainty methods often rely on answer agreement, sample consistency, or entropy, which describe output variation but do not model the numerical meaning of self-reported confidence. Conversely, direct averaging or heuristic aggregation of elicited confidence cannot learn prompt- and task-dependent bias. We propose DirEAG, a Dirichlet Evidence Aggregation method that converts each elicited answer-confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct. Experiments on GSM8K, SVAMP, and GSM-Hard with Qwen, Mistral, and Gemma models show that, compared with direct confidence averaging and heuristic confidence-steering aggregation, DirEAG often achieves better calibration while maintaining competitive answer selection. Ablations further reveal that evidence aggregation and final binary calibration address distinct parts of the calibration problem.

Comments16 pages, 3 figures. Code is available at https://github.com/horacehsugithub/DirEAG. Accepted by PRICAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑