Re:CAP - 审计生产环境RAG管道中的检索覆盖率
Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines
- Machine Learning Center of Excellence, JPMorgan Chase & Co.(摩根大通机器学习卓越中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对生产RAG检索质量难以监控的问题,提出无参考审计方法Re:CAP,通过迭代探测缺失文档,在多个基准上显著提升召回率,且结果稳定可复现。
AI中文摘要:
检索增强生成(RAG)在生产环境中难以监控:对于实时重新索引的非平稳、数百万段落规模的语料库,不存在详尽的相关性标签。因此,检索质量通常研究不足,并且常常被降级优先处理,转而关注面向生成的指标。在这项工作中,我们提出通过探测缺失文档的证据来审计检索覆盖率,而不是枚举所有相关文档。我们的方法Re:CAP(通过迭代探测进行检索覆盖率审计)是一种无参考的审计循环,应用于已部署RAG管道的初始答案和检索上下文:它识别已覆盖的主题,为可能缺失的主题生成探测问题,检索候选文档,并应用LLM作为评判者,仅保留那些引入先前未检索信息的文档。在四个公开基准上,Re:CAP恢复了平面BM25 top-500无法达到的金标标签的9-29%,在TREC-COVID上上升到48%。在MuSiQue上,Re:CAP在不到一半文档预算的情况下,召回率比平面混合top-500高出12.9个百分点。一个集成的BM25、稠密和混合基线(每个top-500)在TREC-COVID上仍然遗漏了21.2%的金标文档,而Re:CAP恢复了这些文档;人工标注者判断,这些结构上不同的文档中有78.9%为基线答案添加了新信息(Fleiss κ = 0.79,n = 123),在实时生产流量中为73.9%(n = 180)。端到端召回率在三次独立运行中可重复性在±1%以内,使Re:CAP成为定期检索审计的稳定工具。
英文摘要:
Retrieval-augmented generation (RAG) is hard to monitor in production: exhaustive relevance labels do not exist for non-stationary multi-million-passage corpora that re-index in real time. As a result, retrieval quality is generally understudied and often deprioritised in favour of generation-oriented metrics. In this work, we propose auditing retrieval coverage by probing for evidence of missing documents rather than enumerating every relevant one. Our method Re:CAP (REtrieval Coverage Audit by iterative Probing) is a reference-free audit loop applied to a deployed RAG pipeline's initial answer and retrieved context: it identifies the topics already covered, generates probing questions for plausibly missing topics, retrieves candidate documents, and applies an LLM-as-judge to retain only those that introduce previously-unretrieved information. On four public benchmarks, Re:CAP recovers 9-29% of gold labels that flat BM25 top-500 cannot reach, rising to 48% on TREC-COVID. On MuSiQue Re:CAP beats flat hybrid top-500 by +12.9 pp on recall at less than half the document budget. An ensemble BM25, dense, and hybrid baseline (top-500 each) still leaves out 21.2% of gold docs on TREC-COVID that Re:CAP recovers; human annotators judge that 78.9% of those structurally distinct documents add new information to the baseline answer (Fleiss $κ$ = 0.79, n = 123), and 73.9% on live production traffic (n = 180). End-to-end recall is reproducible to within $\pm$1% across three independent runs, making Re:CAP a stable instrument for periodic retrieval audits.