arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15145cs.CL

通过推理过程错误分类提升大型语言模型的数学推理能力

Improving Mathematical Reasoning Capabilities in Large Language Models via Reasoning Process Error Classification

  • Human Informatics Labs., NTT, Inc.(NTT公司人类信息学实验室)

机构由 AI 辅助整理,请以论文原文为准。

Runa Yoshida, Kosuke Nishida, Kyosuke Nishida

AI总结:

本研究通过定义并分类21种推理错误,设计聚焦八种错误类别的提示词,有效提升了大语言模型在数学推理任务上的性能。

AI中文摘要:

大型语言模型(LLMs)的推理能力是基于LLM的实际应用中的关键因素。为了探究LLMs当前的推理能力,我们澄清了LLMs在数学数据集上的推理过程中出现的错误类型。我们重点关注LLMs产生错误答案的问题。我们将推理过程中的错误定义为推理错误,并手动分析推理错误的特征。我们定义并分类了21种错误类别,并识别出其中频繁出现的类别。除了定性评估外,我们利用评估结果来提升推理能力。我们设计了一个明确关注八种错误类别的提示词。实验表明,该提示词有效提升了推理性能。此外,结果表明,本文识别出的频繁推理错误在规模相当的LLMs中是普遍存在的。

英文摘要:

The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications. To investigate the current reasoning capability of LLMs, we clarify the types of errors that arise in LLMs' reasoning processes on mathematical datasets. We focus on problems where LLMs produce an incorrect answer. We define errors in the reasoning process as reasoning errors and manually analyze the features of reasoning errors. We defined and classified 21 error classes and identified the frequently occurring classes among them. Beyond qualitative evaluation, we leverage the evaluation results to improve the reasoning capability. We designed a prompt that explicitly focuses on eight error classes. The experiments demonstrate that this prompt effectively improves reasoning performance. Furthermore, the results suggest that the frequent reasoning errors identified in this paper are common across LLMs of comparable scale.

补充信息

↑