arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RetrievalRouter:面向文档检索的联合模态与架构选择

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval

Emre Kuru, Mehmet Onur Keskin, Reza Farahbakhsh, Noel Crespi

arXiv 2608.25625首次发表:更新:

AI 中文总结

RetrievalRouter是一种轻量级查询感知路由模型,仅从查询文本为每个文档检索任务匹配合适流程,在金融、科学语料库基准测试中,其准确率较最佳静态基准高2.5%、速度快12.4倍,且优于现有自适应策略选择方法。

AI 中文摘要

文档检索越来越多地为金融、医疗和法律领域的高风险信息访问提供支持。现代检索流程在模态(文本或多模态)和检索架构(密集型或晚期交互型)两方面存在差异,这些选择带来了难以兼顾的权衡:最有效的流程运行速度过慢且成本过高,无法大规模部署,而最快的流程则无法从复杂文档中检索证据。因此,从业者必须在遗漏证据和无法使用的延迟之间做出选择,且无法在查询层面灵活调整选择依据。我们证明这种权衡并非必要,并非每个查询都需要相同的流程。在涵盖金融和科学语料库的多个基准测试中,没有任何静态流程占据主导地位。我们提出RetrievalRouter,这是一种轻量级的查询感知路由模型,仅从查询文本中学习,为每个查询匹配合适的检索流程。一个可调参数可展示完整的准确率-延迟权衡前沿,对于每个静态基准,RetrievalRouter都提供了一个同时更准确且更快的操作点。与最佳静态基准相比,RetrievalRouter的准确率提升了2.5%,速度提升了12.4倍。此外,与现有自适应策略选择方法相比,RetrievalRouter在面向准确率的设置中实现了显著更高的nDCG@5,在面向延迟的设置中,在nDCG@5和延迟方面均达到或在数值上优于现有方法。我们的代码和数据可在该https链接获取。

英文摘要

Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retrieval pipelines vary both in modality (text or multimodal) and in retrieval architecture (dense or late-interaction). These choices impose a hard compromise: the most effective pipelines are too slow and expensive to run at scale, while the fastest fail to retrieve evidence from complex documents. Practitioners must therefore choose between missed evidence and unusable latency, with no principled basis for adapting that choice at the query level. We show that this compromise is unnecessary. Not every query requires the same pipeline. Across benchmarks spanning financial and scientific corpora, no static pipeline dominates. We introduce RetrievalRouter, a lightweight query-aware router that learns, from the query text alone, which retrieval pipeline best fits each query. A single tunable parameter exposes the full accuracy-latency frontier, and for every static baseline, RetrievalRouter offers an operating point that is simultaneously more accurate and faster. Against the best static baseline, RetrievalRouter is 2.5% more accurate and 12.4 times faster. Furthermore, compared with prior adaptive strategy selection methods, RetrievalRouter achieves significantly higher nDCG@5 across accuracy-oriented settings, while matching or numerically outperforming them on both nDCG@5 and latency in latency-oriented settings. Our code and data are available at https://github.com/emrekuruu/retrieval-router.

CommentsAccepted at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑