发表机构
NVIDIA; University of Edinburgh(英伟达; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出智能体检索,结合LLM推理与检索器,在ReAct循环中解决复杂任务,较标准检索nDCG@10提升8.7点,但耗时和令牌消耗显著增加。
AI 中文摘要
现代信息系统,包括许多智能体工作流,使用稠密检索来探索大量非结构化数据。然而,稠密检索依赖于表面层面的语义相似性,这对于日益复杂的搜索应用来说是不够的。在此,我们研究了智能体检索,它将大型语言模型(LLMs)的推理能力与检索器在ReAct智能体循环中的高效语料库探索相结合,以解决复杂的检索任务。在我们的实验中,我们表明智能体检索比标准检索更有效,使用相同的嵌入模型,nDCG@10提高了8.7个百分点。此外,虽然专门的检索方法在域外任务上表现不佳,但智能体检索具有高度的泛化性:同一流程在ViDoRe v3和BRIGHT排行榜上均取得了有竞争力的结果。然而,这种改进是有代价的。平均而言,智能体检索需要107.4秒,而标准检索只需0.67秒,并且每次查询消耗764.1K个输入和5.8K个输出令牌。总之,我们的研究证明了智能体检索在现代数据系统中的有效性,并激励了未来在大规模部署中开发更具成本效益的检索代理的工作。
英文摘要
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.
CommentsCode: https://github.com/NVIDIA/NeMo-Retriever/tree/main/retrieval-bench