arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.15050cs.AIcs.CL

不要犹豫!过度推理损害置信度校准

Don't Think Twice! Over-Reasoning Impairs Confidence Calibration

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Romain Lacombe, Kerrie Wu, Eddie Dilworth

更新

AI总结:

本文通过 ClimateX 扩展数据集评估推理预算对置信度校准的影响,发现增加推理预算反而导致过度自信,而搜索增强生成以 89.3% 的准确率显著优于纯推理,表明信息获取才是关键瓶颈。

AI中文摘要:

作为问答工具部署的大语言模型需要稳健的校准以避免过度自信。我们使用 ClimateX 数据集(Lacombe et al., 2023)并将其扩展到人类与地球健康领域,系统评估了推理能力和预算如何影响置信度评估的准确性。我们的关键发现挑战了“测试时扩展”范式:虽然近期的推理型大语言模型在评估专家置信度方面达到了 48.7% 的准确率,但增加推理预算始终损害而非改善校准。延长推理会导致系统性过度自信,且随着思考预算的增加而恶化,在适度的计算投入之后产生递减甚至负面的回报。相反,搜索增强生成大幅优于纯推理,通过检索相关证据达到了 89.3% 的准确率。我们的结果表明,对于知识密集型任务的置信度校准改进而言,信息获取而非推理深度或推理预算,可能是关键的瓶颈。

英文摘要:

Large Language Models deployed as question answering tools require robust calibration to avoid overconfidence. We systematically evaluate how reasoning capabilities and budget affect confidence assessment accuracy, using the ClimateX dataset (Lacombe et al., 2023) and expanding it to human and planetary health. Our key finding challenges the "test-time scaling" paradigm: while recent reasoning LLMs achieve 48.7% accuracy in assessing expert confidence, increasing reasoning budgets consistently impairs rather than improves calibration. Extended reasoning leads to systematic overconfidence that worsens with longer thinking budgets, producing diminishing and negative returns beyond modest computational investments. Conversely, search-augmented generation dramatically outperforms pure reasoning, achieving 89.3% accuracy by retrieving relevant evidence. Our results suggest that information access, rather than reasoning depth or inference budget, may be the critical bottleneck for improved confidence calibration of knowledge-intensive tasks.

补充信息

↑