arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScholarCatalyst:一个用于检索激发新研究论文的基准

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn

arXiv 2610.02202首次发表:更新:

发表机构

Stanford University; Seoul National University; University of Washington; Carnegie Mellon University; Allen Institute for AI; MIT(斯坦福大学; 首尔国立大学; 华盛顿大学; 卡内基梅隆大学; 艾伦人工智能研究所; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该基准由作者标注构建,评估检索激发新研究的论文,发现智能体搜索不优于嵌入检索,需新训练方法以赋予模型专家检索直觉。

AI 中文摘要

是什么让伟大的科学家变得伟大?即使人工智能系统开始在开放问题上取得进展,科学家在感知新问题需要哪个先前想法(这些想法埋藏在不断增长的研究档案中)方面仍然远远领先于它们。为了研究这项技能,我们借鉴了那些亲身了解哪些早期工作推进了他们已完成项目的研究人员,论文作为指向其中思想的指针。利用我们使作者注释可扩展的自动化流程,我们通过让207篇近期计算机科学论文的184位主要作者标记哪些候选论文已经或可能推进了他们的项目(每个都附有详细理由),构建了ScholarCatalyst。我们引入了一个带有作者提供判断的检索任务:给定一个初始研究问题,仅从项目开始时可用的文献中检索这些论文。尽管智能体搜索调用相同的检索器作为工具,但其表现并不优于嵌入检索(Recall@20分别为0.42和0.48)。即使是基于Claude Fable 5.1构建的智能体(可能在训练期间见过已完成的论文),其R@20也仅达到0.51。这些结果凸显了对新训练配方的需求,这些配方能使模型具备搜索广泛语料库的专家直觉。我们将ScholarCatalyst视为迈向科学智能体的一步,这些智能体能够将一个半成形的想法指向其所需的先前研究。

英文摘要

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

Comments57 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑