面向大语言模型复杂图推理的统一多维度基准
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
浏览论文内容
中文总结 AI 辅助
针对现有图推理基准的局限,提出半自动框架构建含202个任务的统一多维度基准,评估LLM在不同推理模式下的表现,揭示模型局限并提供实证指导。
中文摘要 AI 辅助
图推理为评估大语言模型(LLM)的推理能力提供了极具潜力的测试平台,因为图实例可通过编程生成、结构可控,且能自然扩展至长输入场景。然而,现有图推理基准对数据复杂度的覆盖有限,过度依赖人工构建,且缺乏对文本式与代码式推理模式的统一评估。为解决这些局限,我们提出了{\texttt{GraphReasonBench}}(图推理基准),这是一个用于构建复杂图推理基准的五阶段半自动框架。该框架沿五个维度拓展基准覆盖范围:图规模、任务复杂度、任务描述、图加载及任务来源。该框架采用基于LLM的数据生成器自动生成任务描述、图数据、参考解决方案、图加载脚本、问题形式及评估脚本,同时在关键质量控制阶段保留人工验证。基于此,我们构建了包含202个任务的基准,并在文本式、代码式及增强式推理设置下评估LLM。实验表明,这些复杂度维度揭示了现有基准中不易察觉的模型局限;现有微调模型难以泛化至GraphGym,而检索增强方法呈现出场景依赖的适应性,可提升文本推理能力但无法始终改善编码推理能力。这些发现表明,我们的基准可作为具有挑战性和诊断性的图推理基准,并为未来增强方法提供实证指导。代码与数据集即将发布。
英文摘要
Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings. However, existing graph reasoning benchmarks have limited coverage of data complexity, rely heavily on manual construction, and lack unified evaluation across text-based and code-based reasoning modes. To address these limitations, we propose {\dataset}, a five-stage \textit{semi-automatic} framework for constructing complex graph reasoning benchmarks. It expands benchmark coverage along five dimensions: \textit{Graph Size}, \textit{Task Complexity}, \textit{Task Description}, \textit{Graph Loading}, and \textit{Task Source}. The framework uses an LLM-based data generator to automatically produce task descriptions, graph data, reference solutions, graph-loading scripts, question forms, and evaluation scripts, while retaining human validation at key quality-control stages. Based on it, we construct a benchmark with $202$ tasks and evaluate LLMs under text-based, code-based, and augmented reasoning settings. Experiments show that the complexity dimensions reveal model limitations that are less visible in existing benchmarks; existing fine-tuned models struggle to generalize to GraphGym, whereas retrieval-augmented methods show scenario-dependent adaptability, improving textual reasoning but not consistently improving coding reasoning. These findings suggest that ours serves as a challenging and diagnostic benchmark for graph reasoning and provides empirical guidance for future enhancement methods. Code and dataset will be published soon.
发表机构
- The Pennsylvania State University(宾夕法尼亚州立大学)
- Rensselaer Polytechnic Institute(伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。