arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMBL AI Librarian:面向AI智能体的生命科学知识层

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Luigi Sigillo, Matteo Silvestri, Francesco Tabaro, Rajat Bhatnagar, Syed Irtaza Mubashar, Matt Jeffryes, Daljit Nijjer, Vittorio Perera, Ola Spjuth, Julio Saez-Rodriguez, Melissa Harrison, Fabio Petroni

arXiv 2607.28229首次发表:更新:

AI 中文总结

该研究推出EMBL AI Librarian,为AI智能体升级Europe PMC接口,通过LLM统筹检索,在多基准任务中提升性能,代码已公开。

AI 中文摘要

网络正越来越多地被AI智能体而非人类访问。每一个智能体都需要知识,尤其是在生命科学领域,智能体流程发展迅速。获取文献是这一需求的关键部分,拥有超过4000万条索引记录的Europe PMC等资源被广泛用于满足这一需求。然而这些资源并非为AI智能体而构建:它们接收关键词和复杂语法并返回完整论文,因此每个智能体必须学习语法、发起多次搜索并阅读完整论文以找到所需证据。我们推出EMBL AI Librarian,这是一个为AI智能体升级Europe PMC接口的知识层:智能体用自然语言提问,即可获得能回答问题的证据。单个大语言模型(LLM)统筹整个知识检索过程:它规划由实时Europe PMC搜索引擎执行的互补子查询,随后读取选定的论文并定位相关证据。我们在四个基准上评估Librarian:文献合成、声明验证、开放域问答,以及下游生物学任务,如方案问题和序列操作。在ScholarQABench上,Librarian相比近期发布的强大基线将Citation F1提升了16个百分点以上;作为现有声明验证流程的检索层使用时,它提高了与专家共识的一致性;在开放形式的LitQA2基准上,基于Librarian的GPT-5.4智能体比使用网络搜索时得分高出约8个百分点。总体而言,我们的结果表明,为生命科学智能体配备Librarian知识层可提升一系列任务的性能。我们将代码公开发布在此处https URL。

英文摘要

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than $16$ points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about $8$ points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑