arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.07820cs.AIcs.CL

在大语言模型中的确定性引导推理:一种动态思维预算方法

Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach

  • Nokia Bell Labs, France(诺基亚贝尔实验室,法国)

机构由 AI 辅助整理,请以论文原文为准。

João Paulo Nogueira, Wentao Sun, Alonso Silva, Laith Zumot

更新

AI总结:

CGR通过动态调整推理预算,在保持准确性的同时减少token使用,通过确定性评估实现高效的推理过程。

AI中文摘要:

大型推理语言模型通常使用固定的推理预算,这可能导致计算浪费或提前终止推理。我们引入了确定性引导推理(CGR),这是一种模型无关的自适应推理过程,定期探测当前推理是否支持有信心的最终答案,并在达到目标确定性阈值时提前终止,否则继续直到结束思考的标记或预算限制。确定性是从模型对答案标记的预测概率中估计的,从而得到一个轻量的停止标准。在AIME2025上,CGR保持基线准确性的同时减少token使用,提供了一个可调节的确定性-效率权衡,可以消除数百万个token。在64个随机种子中,CGR表现出一致的行为。我们还引入了一个Grade指标,惩罚错误答案并允许弃权,捕捉风险敏感的性能。结果表明,CGR在确定性较低时通过弃权来提高Grade。

英文摘要:

Large reasoning language models are typically run with fixed inference budgets, which can waste computation or terminate reasoning prematurely. We introduce Certainty-Guided Reasoning (CGR), a model-agnostic adaptive inference procedure that periodically probes whether the current reasoning supports a confident final answer and terminates early once a target certainty threshold is reached, otherwise continuing until the end-of-thinking token or the budget limit. Certainty is estimated from the model's predicted probabilities over the answer tokens, yielding a lightweight stopping criterion. On AIME2025, CGR preserves baseline accuracy while reducing token usage, providing a tunable certainty-efficiency trade-off that can eliminate millions of tokens in aggregate. Across 64 random seeds, CGR exhibits consistent behavior. We also introduce a Grade metric that penalizes incorrect answers and permits abstention, capturing risk-sensitive performance. Results show that CGR improves Grade by abstaining when certainty remains low.

↑