arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciExplore:评估从科学导航到信息整合的自主智能体

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

Yinhao Tang, Youqing Fang, Yanan Sun, Wenran Liu, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen

arXiv 2607.20926首次发表:更新:

发表机构

University of Science and Technology of China; Shanghai AI Laboratory(中国科学技术大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍用于评估大语言模型和智能体科学信息检索与推理能力的基准SciExplore,包含四类任务,在其上评估多个先进模型和智能体,发现随着任务复杂度增加性能差距大,凸显当前模型在实际科学信息检索场景中的局限。

AI 中文摘要

科学研究涉及跨异构源的复杂信息检索和推理工作流程。然而,现有基准主要强调通用领域检索或静态科学问答,无法评估实际科学研究工作流程所需的关键能力。我们引入了SciExplore,这是一个旨在评估大语言模型和智能体科学信息检索和推理能力的基准。SciExplore包括四种任务类型,涵盖十多个科学学科的103个专家策划任务:科学数据库导航、模糊文献检索、缺失参考文献补全和跨源结构化知识合成,从实体级推理、文档级识别到证据级基础和领域级合成,逐步探索更高层次的能力。我们在SciExplore上评估了十多个先进的大语言模型和自主智能体,结果显示随着任务复杂性增加,性能差距显著,在最具挑战性的结构化合成任务上准确率极低。这些结果凸显了当前模型和智能体在实际科学信息检索场景中的重大局限性。

英文摘要

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

Comments25 pages, 13 figures. Submitted to ACL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑