arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用Prolog探测语言模型的长程演绎推理能力

Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

Hadeel Al-Negheimish, Jasna Ilieva, Yoon Kim

arXiv 2610.11592首次发表:更新:

发表机构

King Saud University; Massachusetts Institute of Technology(沙特国王大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了Prolog长推理测试平台ProloNg,探究前沿LLM的长程演绎推理能力,发现其性能随推理深度增加大幅下降,深度超10时多数模型接近随机水平。

AI 中文摘要

当前前沿大语言模型(LLM)理论上可处理100万token及以上的长上下文,但它们在多大程度上能超越简单检索,对这类长上下文进行更深入的推理?我们实证研究了LLM的长程推理能力,重点关注用Prolog表达的演绎逻辑。我们构建了ProloNg——一个用于探测Prolog长推理的合成测试平台,该平台会系统地改变问题的复杂度(推理深度),其中最难的案例推理深度达22,上下文长度为6.2万。我们研究了5类前沿LLM中的8个推理模型,发现随着推理深度增加,性能会大幅下降,多数模型在深度超过10时接近随机水平。

英文摘要

Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.

CommentsFindings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑