ScalableRAG:零摄入成本下的高质量检索增强生成
ScalableRAG: High-Quality RAG at Zero Ingestion Cost
浏览论文内容
中文总结 AI 辅助
研究提出零摄入成本的ScalableRAG及有限摄入的改进版本,通过维护工作区实现即时聚合推理,在六个语料库测试中表现出色,平均准确率大幅超越基线,还通过固定LLM调用次数等进一步提升大规模时的准确率。
中文摘要 AI 辅助
检索增强生成(RAG)的最新进展旨在通过为知识摄入支付高昂成本来优化性能,比如构建知识图谱或提取SQL表。本文表明,此类知识库所允许的操作可以零摄入成本复制(甚至无需向量数据库)。实际上,我们的零摄入ScalableRAG解决方案在此处考虑的六个语料库中的三个中轻松超越所有基线(包括知识图谱方法),在另外三个中仅略低于最佳性能,所有六个数据集的平均准确率比第二有竞争力的基线高7.36%。它通过维护可读写的文档集和值集工作区来实现,在需要根据与文档集子集一一对应的主键进行分组的所有情况下都能进行即时聚合推理。我们还引入了有限摄入ScalableRAG,通过固定LLM调用次数且独立于语料库大小,并使用最小向量数据库以及从文档样本中自动发现模式,来进一步提高大规模时的准确率。代码可通过给定链接获取。
英文摘要
Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or extracting SQL tables. In this work we show that the operations that such knowledge bases allow can be replicated with zero ingestion costs (not even a vector database); in fact our solution, Zero-Ingestion ScalableRAG, handily out-performs all baselines (including knowledge graph approaches) in three out of the six corpora considered here, and only marginally missing maximum performance on the other three, with average accuracy across all six datasets 7.36% above the next most competitive baseline. It achieves this by keeping a workspace of document sets and values sets that it can write into and read from, allowing for on-the-fly aggregative reasoning in all situations where grouping is required on a primary key that is in one to one correspondence with a subset of the total document set. Capping the number of LLM calls by a constant independent of the corpus size, we also introduce Limited-Ingestion ScalableRAG, which does use a minimal vector database as well as an automated pattern discovery from a sample of documents, to further improve accuracy at scale. Our code is available at https://github.com/cohesity/ScalableRAG .
发表机构
- Cohesity(科赫思)
机构由 AI 辅助整理,请以论文原文为准。