发表机构
JPMorganChase(摩根大通)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出增量式池化LLM评估方法,通过复用判断降低RAG系统检索模型选择的评估成本,在多基准测试中验证了其有效性,可实现高性价比的增量检索模型选择。
AI 中文摘要
为生产环境中的检索增强生成(RAG)系统选择检索模型需要可靠的对比评估,但获取大规模相关性判断(qrels)成本高昂,且随着新候选系统的加入难以重复开展。本文研究池化大语言模型(LLM)评估方法:由LLM对当前候选系统集合检索到的全部文档进行判断,当引入新系统时,仅对其贡献的新文档进行判断,再复用这些判断结果以统一标准评估所有系统。我们在4个涵盖密集型、稀疏型和混合型配置的检索基准上对该方法进行验证,还将其部署用于对比金融新闻问答系统的62种检索配置。池化LLM排名与各数据集的黄金标准评估高度相关,在考虑qrels的自举不确定性后,97%的系统成对排序得以保留。在生产环境中,文档重叠可实现65%-80%的判断复用,评估成本最高降低4.9倍,使团队无需重新判断已评估文档即可对新检索候选进行基准测试。这些结果表明,池化LLM评估是部署系统中增量式检索模型选择的实用且高性价比的工作流程。
英文摘要
Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtaining relevance judgments at scale is expensive and difficult to repeat as new candidate systems arrive. We study pooled LLM evaluation, in which an LLM judges the union of documents retrieved by the current set of candidate systems, and the pool is then expanded incrementally as new systems are introduced by judging only the new documents they contribute. These judgments are reused to evaluate all systems on a common basis. We validate this approach on four retrieval benchmarks with 11 systems spanning dense, sparse, and hybrid configurations, and deploy it to compare 62 retrieval configurations for a financial news QA system. Pooled LLM rankings correlate strongly with gold-standard evaluation across datasets, and 97% of pairwise system orderings are preserved once bootstrap uncertainty in the qrels is taken into account. In production, document overlap yields 65-80% judgment reuse and up to 4.9x lower evaluation cost, allowing teams to benchmark new retrieval candidates without re-judging previously assessed documents. These results suggest pooled LLM evaluation is a practical and cost-effective workflow for incremental retrieval model selection in deployed systems.
Comments10 pages, 1 figure