arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15624cs.IR

检索器能否从不同方面找到同一篇论文?一个多方面全论文科学检索基准

Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark

Yiyang Wei, Fang Guo, Qiji Zhou, Zhizhang Fu, Mengru Ding, Kai Yang, Yue Zhang

AI总结:

该研究推出多方面全论文科学检索基准MAPLE及查询生成流程MAPLE-Synth,发现现有检索器在从单一方面与多方面检索论文间存在显著差距,为相关检索器研发提供测试平台。

AI中文摘要:

科学论文包含背景、方法等多个可检索方面,但许多论文检索基准仅评估查询与论文的个体相关性,忽略同一论文的其他方面。为填补这一空白,我们推出MAPLE,这是一个经专家验证的多方面全论文检索基准,用于评估检索器能否从针对论文动机、方法和实验结果的查询中一致地找回同一篇论文。MAPLE包含2095个关于近期ML和NLP论文的查询,基于文本和多模态内容。我们进一步提出MAPLE-Synth,这是一种基于检索的上下文学习流程,利用OpenReview讨论和人工编写的查询示例生成反映研究人员对论文不同方面兴趣的真实查询。我们的专家验证表明,这些查询在真实性上可与人工编写的查询相媲美,且与目标论文高度相关。对词汇、科学领域、通用文本和多模态检索器的实验显示,从任意一个方面检索论文与从所有方面检索论文之间存在巨大差距:最强模型在AnyAspect@20指标上达到98.1%,但在AllAspect@20指标上仅为15.7%。跨检索器来看,实验/结果查询和引用表格的查询尤其困难。尽管多块聚合改进了多方面论文检索,但仍存在大量失败案例。MAPLE为评估和开发能更全面表示科学论文的检索器提供了测试平台。

英文摘要:

Scientific papers contain multiple searchable facets such as background, methods. However, many paper retrieval benchmarks merely evaluate individual query-paper relevance, while overlooking other facets of the same paper. To bridge this gap, we introduce MAPLE, an expert-validated benchmark for multi-aspect, full-paper retrieval that evaluates whether retrievers can consistently recover the same paper from queries targeting its motivation, method, and experimental findings. MAPLE contains 2,095 queries about recent ML and NLP papers, grounded in both textual and multimodal content. We further propose MAPLE-Synth, a retrieval-based in-context learning pipeline that leverages OpenReview discussions and human-written query exemplars to generate realistic queries reflecting researchers' interests in different aspects of a paper. Our expert validation shows that these queries are comparable in realism to human-written queries and highly relevant to the target papers. Experiments across lexical, scientific-domain, general-purpose text, and multimodal retrievers reveal a substantial gap between retrieving a paper from any one aspect and retrieving it from all aspects: the strongest model achieves 98.1% AnyAspect@20 but only 15.7% AllAspect@20. Experiment/result queries and table-referenced queries are particularly difficult across retrievers. Although multi-chunk aggregation improves multi-aspect paper retrieval, considerable failures persist. MAPLE provides a testbed for evaluating and developing retrievers that represent scientific papers more comprehensively.

↑