arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20498cs.AI

AISE-Bench:用于学术知识图谱信息检索的全周期精选基准测试

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对学术知识图谱信息检索基准测试不足的问题,引入AISE-Bench,通过定制工作流程和综合评估协议构建基准测试,为多步API使用的大语言模型智能体提供新测试平台,助力评估与改进。

中文摘要 AI 辅助

随着配备工具的大语言模型成为能够利用网络引擎、应用程序编程接口(API)和代码解决复杂长期任务的自主智能体,当前学术图谱信息检索的工具使用基准测试存在不足。我们引入了AISE-Bench,一个用于学术知识图谱信息检索的真实世界、全周期注释基准测试。它包含1133个问答对,设计了定制智能体工作流程以支持高质量注释,还开发了综合评估协议。在14种评估方法中,即使最强模型表现也一般,AISE-Bench为评估和改进多步API使用的大语言模型智能体提供了新测试平台。

英文摘要

Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic graphs rely on synthetic templates, simplified solution spaces, or narrow tasks such as paper-centric tasks, leaving key challenges underexplored - realistic user intent, complex multi-step API planning, rich parameter filling for APIs, grounded answers with references, and comprehensive evaluation of both the process and the outcome. We introduce AISE-Bench, a real-world, full-cycle annotated benchmark for information seeking on academic knowledge graphs. AISE-Bench release contains 1,133 QA pairs, including query taxonomies, full API execution trajectories, validated parameters, and source-grounded answers with reference links. To support high-quality annotation, we design a customized agent workflow to enable annotators to plan, execute, and revise complex API workflows efficiently. We develop a comprehensive evaluation protocol measuring answer quality, reference grounding, API-planning correctness, and execution success. Among the 14 evaluated methods, even the strongest model (PLAY2PROMPT with Gemini-3-Pro) achieves only moderate performance and often struggles with API planning and execution. AISE-Bench establishes a challenging new testbed for quantitatively evaluating and improving the stepwise correctness, grounded summarization, and traceable reasoning of multi-step API-using LLM agents. Our code and data are available at https://aise-bench.github.io/.

发表机构

  • Renmin University of China(中国人民大学)
  • Anhui University(安徽大学)
  • Z-Lab Z.ai
  • Tsinghua University(清华大学)
  • Bosch Center for AI(博世人工智能中心)
  • University of Oslo(奥斯陆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑