arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INSPIRE:面向开放研究问题的科学文献检索基准

Inspire: Benchmarking Scientific Literature Search for Open Research Problems

Jianrong Ding, Zhengyan Shi, Jianyuan Zhong, Kai Qiu, Qi Dai, Yifan Yang, Chong Luo, Qiang Xu

arXiv 2609.33233首次发表:更新:

发表机构

The Chinese University of Hong Kong; Microsoft Research Asia(香港中文大学; 微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

INSPIRE是一个面向开放研究问题的科学文献检索基准,通过分阶段诊断揭示资源暴露是主要瓶颈,最强智能体nDCG@10仅为0.284,并验证了可重放演示能提升检索性能。

AI 中文摘要

科学文献检索通常始于一个开放的研究问题,而非已知的目标论文或固定的候选集。我们引入了INSPIRE,一个用于评估智能体在检索先前文献以推进解决方案被隐去的研究问题方面的基准。每个实例将一份研究简报与一个目标特定的截止时间配对,该截止时间设定在后续论文发表前三个月,并根据该论文实际研究谱系中分级的被引先例来评估排序输出。检索在开放语料库上进行,而目标论文的身份及其被引先例的成员资格对智能体保持隐藏。除了端到端的检索质量外,INSPIRE利用记录的检索轨迹来区分三个耦合阶段:资源暴露,即检索过程中是否有用的先例被呈现;选择,即暴露的先例是否被保留;以及排序,即保留的论文被排序的有效程度。在共享检索界面和预算下,针对476个计算机科学目标,评估的最强智能体达到了0.284的nDCG@10。结果表明,当前智能体更容易恢复单个孤立的先例,而不是整合更广泛的相关先前工作组合。分阶段分析确定资源暴露是观察到的最大瓶颈,其次是选择和排序中的进一步损失。我们还构建了可重放的追溯演示,并表明它们在不改变测试时信息的情况下改善了保留集上的搜索,从而确立了该基准提供了可操作的学习信号。因此,INSPIRE在智能体必须构建自身相关工作标准的场景中,既支持端到端比较,也支持分阶段诊断。

英文摘要

Scientific literature search often begins with an open research problem rather than a known target paper or a fixed candidate set. We introduce INSPIRE, a benchmark for evaluating agents that search prior literature to make progress on solution-redacted research problems. Each instance pairs a research brief with a target-specific cutoff three months before a later paper and evaluates ranked outputs against graded cited antecedents from that paper's realized research lineage. Search proceeds over an open corpus, while the identity of the target paper and membership of its cited antecedents remain hidden from the agent. Beyond end-to-end retrieval quality, INSPIRE uses logged search trajectories to distinguish three coupled stages: resource exposure, whether useful antecedents are surfaced during search; selection, whether exposed antecedents are retained; and ranking, how effectively retained papers are ordered. Across 476 computer-science targets under a shared search interface and budget, the strongest evaluated agent achieves 0.284 nDCG@10. Results show that current agents more readily recover an isolated antecedent than assemble a broader portfolio of relevant prior work. The stagewise analysis identifies resource exposure as the largest observed bottleneck, with further losses in selection and ranking. We additionally construct replay-valid hindsight demonstrations and show that they improve held-out search without changing test-time information, establishing that the benchmark provides an actionable learning signal. INSPIRE therefore enables both end-to-end comparison and stage-resolved diagnosis in a setting where the agent must construct its own working criterion of relevance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑