文档检索器的系统性多域评估
A Systematic Multi-Domain Evaluation of Document Retrievers
浏览论文内容
中文总结 AI 辅助
本研究对33个文档检索器在七个数据集上进行大规模实证评估,发现NV-Embed-v2性能最强但延迟高,SPLADE-v3以低延迟媲美顶尖方法,并揭示了检索器失败点中的潜在改进空间。
中文摘要 AI 辅助
文档检索是现代许多人工智能系统的关键组成部分,直接影响其在下游任务中的有效性、鲁棒性和公平性。尽管近年来检索器的数量不断增长,但文献中的比较研究通常在范围上有限,或聚焦于单一基准、领域或模型家族。这种碎片化使得难以就文档检索器的相对优势、劣势和权衡得出可靠结论。为解决这一差距,我们对文档检索器进行了大规模实证评估,涵盖三个家族(稀疏、稠密和基于扩展的),并在七个信息检索数据集上评估了33个检索器,分析了检索质量、运行时间和失败点。我们没有对每个模型进行单独调优,而是对每个检索器进行开箱即用的评估,使用可从其公开文档重建的配置和统一的计算预算。我们的结果显示,NV-Embed-v2在七个数据集中的四个上取得了最强性能,尽管以显著的查询延迟为代价。在稀疏检索器中,我们发现SPLADE-v3以低得多的延迟媲美顶尖方法,甚至在MS MARCO上取得了最高分。在指令遵循数据集上,GritLM提供了最佳性能。最后,对检索器失败点的分析揭示了模型与家族之间的对比,表明检索性能存在未实现的潜在提升空间。
英文摘要
Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers' failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.
发表机构
- University of Konstanz(康斯坦茨大学)
机构由 AI 辅助整理,请以论文原文为准。