GABench:用于评估大语言模型智能体图分析任务的综合基准
GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
浏览论文内容
中文总结 AI 辅助
本文提出GABench这一综合基准,涵盖多类图类型与任务及84个可执行工具,构建10400个带可验证答案的任务,评估前沿LLM及智能体框架,发现现有智能体应对复杂图任务仍存局限等结论。
中文摘要 AI 辅助
大语言模型(LLM)智能体在规划、使用工具及与外部环境交互方面的能力日益增强,通常由管理状态、协调多步执行的框架提供支持。图分析为评估其智能体能力提供了理想场景,因为该任务要求智能体在图环境中访问数据并执行操作。然而,现有的LLM图基准对图任务和图类型的覆盖有限,难以全面评估LLM智能体;且这些基准通常将图分析表述为基于文本的问答,图信息直接嵌入提示中,限制了对端到端智能体能力的评估。为解决这些局限,本文提出GABench——一个用于智能体图分析的综合基准。GABench涵盖3种图类型,包含4类图分析任务:图检索、图论、图机器学习及图开放式问答;还提供84个可执行工具,用于访问图数据及执行各类图操作。基于这些工具,本文开发了智能体图分析任务生成流程,构建了10400个带有可验证标准答案的任务。利用GABench,本文评估了一系列前沿LLM及智能体框架,实验揭示三个关键发现:1.现有LLM智能体仍难以应对复杂图分析任务;2.框架选择对性能影响显著,但现有框架在复杂图任务上仍存在局限;3.图分析对工具调用质量的依赖高于数量。这些发现为开发和评估用于图分析的LLM智能体提供了实用见解。
英文摘要
Large language model (LLM) agents are increasingly capable of planning, using tools, and interacting with external environments. They are typically supported by harnesses, which manage state and coordinate multi-step execution. Graph analysis provides a promising setting for evaluating their agentic capabilities, because it requires agents to access data and execute operations in a graph environment. However, existing graph benchmarks for LLMs provide limited coverage of graph tasks and graph types, making it difficult to comprehensively evaluate LLM agents. Moreover, they typically formulate graph analysis as text-based question answering, where graph information is directly provided in the prompt, limiting the evaluation of end-to-end agentic capabilities. To address these limitations, we introduce GABench, a comprehensive benchmark for agentic graph analysis. GABench spans three graph types and covers four graph analysis task categories: graph retrieval, graph theory, graph machine learning, and graph open-ended question answering. GABench also provides 84 executable tools for accessing graph data and performing diverse graph operations. Building on these tools, we develop an agentic graph analysis task generation pipeline and construct 10,400 tasks with verifiable ground truth.Using GABench, we evaluate a range of frontier LLMs and agent harnesses. Our experiments reveal three key findings: (1) Existing LLM agents still struggle with complex graph analysis tasks. (2) Harness choice significantly affects performance, yet existing harnesses remain limited on complex graph tasks. (3) Graph analysis depends more on tool-call quality than quantity. Our findings provide practical insights into the development and evaluation of LLM agents for graph analysis.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
- The Chinese University of Hong Kong(香港中文大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。