发表机构
University of Massachusetts, Amherst; VA Bedford Health Care; University of Massachusetts, Lowell(马萨诸塞大学阿默斯特分校; 退伍军人事务部贝德福德医疗保健中心; 马萨诸塞大学洛厄尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出前瞻性知识根源检索基准MUSES及配套工具CiteRoots,实验显示基于SPECTER2的多质心检索器表现最优,为相关领域提供了新的测试平台与方法参考。
AI 中文摘要
科学发现依赖于寻找能塑造后续研究的现有文献。现有检索系统优化相关性与流行度,往往更倾向于核心论文,而非后来被证明具有生成性的较不熟悉的作品。我们推出MUSES,这是一个基于233万篇论文固定语料库的百万级前瞻性知识根源检索基准,每个熟悉度层级约有14万测试实例。据我们所知,它是首个达到此规模的前瞻性基准,拥有统一检索任务和作者确认的论文级根源标签。与此同时,CiteRoots结合了基于局部引用文本的可扩展修辞层(LLM评判者κ=0.896,与人类黄金标准相比)和论文级作者认可层(来自753篇焦点论文的1518个生成性灵感对)。MUSES沿两个维度划分难度:熟悉度维度涵盖CiteNext、CiteNew和CiteNew-Isolated,功能维度涵盖广泛引用、修辞根源和作者认可根源。在9种方法类别中,基于SPECTER2构建的精简多质心检索器表现最佳。Hit@100从CiteNext的0.534降至CiteNew的0.424、修辞CiteNew的0.205以及作者认可CiteNew的0.171,降幅达3.1倍。在注册的8视角全测试审计中,K=1000时,约一半的广泛层级测试实例仍未被解决。修辞角色与作者认可存在差异:同一评判者与认可的一致性为κ=0.037。我们发布MUSES、两个CiteRoots层以及一个蒸馏的开放配套评判者,用于未来前瞻性检索和知识根源的研究。
英文摘要
Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce \textbf{MUSES}, a million-instance benchmark for prospective intellectual-roots retrieval over a fixed 2.33M-paper corpus, with roughly 140K test instances per familiarity tier. To our knowledge, it is the first prospective benchmark at this scale with a shared retrieval task and author-confirmed paper-level root labels. Alongside it, \textbf{CiteRoots} pairs a scalable rhetorical layer over local citation text (LLM judge $κ= 0.896$ versus human gold) with a paper-level author-endorsed layer ($n = 1{,}518$ generative-inspiration pairs from 753 focal papers). MUSES organizes difficulty along two axes: a \emph{familiarity} axis spanning CiteNext, CiteNew, and CiteNew-Isolated, and a \emph{functional} axis spanning broad citations, rhetorical roots, and author-endorsed roots. Across 9 method classes, a lean multi-centroid retriever built on SPECTER2 is strongest. Hit@100 falls from 0.534 on CiteNext to 0.424 on CiteNew, 0.205 on rhetorical CiteNew, and 0.171 on author-endorsed CiteNew, a $3.1\times$ decline. In a registered eight-lens full-test audit, roughly half of broad-tier test instances remain unsolved at K=1{,}000. Rhetorical role and author endorsement are distinct: the same judge agrees with endorsement at $κ= 0.037$. We release MUSES, both CiteRoots layers, and a distilled open companion judge for future work on prospective retrieval and intellectual roots.