LLM 能否进行结构性推理?通过数据结构视角进行基准测试
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 DSR-Bench 基准,通过 20 种数据结构、35 种操作和 4140 个问题实例评估 LLM 的结构性推理能力,发现顶级模型在挑战性实例上仅得 0.46/1,且在空间数据、上下文丰富场景及自身代码推理上表现不佳。
AI中文摘要:
大型语言模型(LLM)被部署在日益复杂的任务上,这些任务需要多步决策。因此,理解它们的算法推理能力至关重要。然而,我们缺乏用于评估这些能力的诊断基准。我们提议使用数据结构作为原则性视角:作为算法的基本构建块,它们自然地探测结构性推理——即理解和操作支撑算法推理的关系(如顺序、层次和连接性)的能力。我们引入了 DSR-Bench(数据结构推理基准),涵盖 20 种数据结构、35 种操作和 4140 个问题实例。DSR-Bench 具有层次化任务组织、全自动生成与评估以及细粒度诊断的特点。评估 13 个最先进的 LLM 揭示了关键局限性:表现最好的模型在挑战性实例上仅达到 0.46/1。三个针对更现实用法的辅助探针暴露了进一步的弱点:模型在空间数据和上下文丰富的场景中表现不佳,并且难以对其自身代码进行推理。
英文摘要:
Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithmic reasoning abilities is therefore crucial. However, we lack a diagnostic benchmark for evaluating these capabilities. We propose to use data structures as a principled lens: as fundamental building blocks of algorithms, they naturally probe structural reasoning - the ability to understand and manipulate relationships such as order, hierarchy, and connectivity that underpin algorithmic reasoning. We introduce DSR-Bench (Data Structure Reasoning Benchmark), spanning 20 data structures, 35 operations, and 4,140 problem instances. DSR-Bench features hierarchical task organization, fully automated generation and evaluation, and fine-grained diagnostics. Evaluating 13 state-of-the-art LLMs reveals critical limitations: the top-performing model achieves only 0.46/1 on challenging instances. Three auxiliary probes targeting more realistic usages expose further weaknesses: models perform poorly on spatial data and context-rich scenarios, and they struggle to reason over their own code.