图表有助于大语言模型推理吗?来自三段论推理的证据
Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning
浏览论文内容
中文总结 AI 辅助
研究三段论推理中图表对大语言模型的影响,比较自然语言、逻辑符号、线性图表和欧拉图四种表示条件,利用285个问题评估两个模型,发现图表不能始终提升性能,模型在中性问题上表现不佳,从图表中获益有限。
中文摘要 AI 辅助
图表被广泛用于支持逻辑推理,先前研究表明像欧拉图这样的表示法可提高人类推理表现。近期工作也探索了其对大语言模型的影响。本文比较三段论推理的四种表示条件:自然语言、逻辑符号、线性图表和欧拉图。利用安藤等人(2024)的285个问题,评估了两个当代大语言模型Claude 3.5~Sonnet和GPT-4o-mini。结果表明图表表示法并非总能提升性能。模型在蕴含和矛盾问题上表现良好,但在中性问题上存在困难且常犯系统性转换错误。总体而言,测试模型在逻辑推理任务中从图表获得的益处有限。
英文摘要
Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four representational conditions for syllogistic reasoning: natural language, logical notation, linear diagrams, and Euler diagrams. Using 285 problems from Ando et al. (2024), we evaluate two contemporary LLMs, Claude 3.5~Sonnet and GPT-4o-mini. Our results show that diagrammatic representations do not consistently improve performance. Although the models perform well on entailment and contradiction problems, they struggle with neutral problems and often make systematic conversion errors. Overall, the results suggest that the tested models gain limited benefit from diagrams in logical reasoning tasks.
发表机构
- Keio University(庆应义塾大学)
机构由 AI 辅助整理,请以论文原文为准。