AISE-Bench:用于学术知识图谱信息检索的全周期精选基准测试
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
浏览论文内容
中文总结 AI 辅助
研究针对学术知识图谱信息检索基准测试不足的问题,引入AISE-Bench,通过定制工作流程和综合评估协议构建基准测试,为多步API使用的大语言模型智能体提供新测试平台,助力评估与改进。
中文摘要 AI 辅助
随着配备工具的大语言模型成为能够利用网络引擎、应用程序编程接口(API)和代码解决复杂长期任务的自主智能体,当前学术图谱信息检索的工具使用基准测试存在不足。我们引入了AISE-Bench,一个用于学术知识图谱信息检索的真实世界、全周期注释基准测试。它包含1133个问答对,设计了定制智能体工作流程以支持高质量注释,还开发了综合评估协议。在14种评估方法中,即使最强模型表现也一般,AISE-Bench为评估和改进多步API使用的大语言模型智能体提供了新测试平台。
英文摘要
Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic graphs rely on synthetic templates, simplified solution spaces, or narrow tasks such as paper-centric tasks, leaving key challenges underexplored - realistic user intent, complex multi-step API planning, rich parameter filling for APIs, grounded answers with references, and comprehensive evaluation of both the process and the outcome. We introduce AISE-Bench, a real-world, full-cycle annotated benchmark for information seeking on academic knowledge graphs. AISE-Bench release contains 1,133 QA pairs, including query taxonomies, full API execution trajectories, validated parameters, and source-grounded answers with reference links. To support high-quality annotation, we design a customized agent workflow to enable annotators to plan, execute, and revise complex API workflows efficiently. We develop a comprehensive evaluation protocol measuring answer quality, reference grounding, API-planning correctness, and execution success. Among the 14 evaluated methods, even the strongest model (PLAY2PROMPT with Gemini-3-Pro) achieves only moderate performance and often struggles with API planning and execution. AISE-Bench establishes a challenging new testbed for quantitatively evaluating and improving the stepwise correctness, grounded summarization, and traceable reasoning of multi-step API-using LLM agents. Our code and data are available at https://aise-bench.github.io/.
发表机构
- Renmin University of China(中国人民大学)
- Anhui University(安徽大学)
- Z-Lab Z.ai
- Tsinghua University(清华大学)
- Bosch Center for AI(博世人工智能中心)
- University of Oslo(奥斯陆大学)
机构由 AI 辅助整理,请以论文原文为准。