TurboVec:通过码本无关量化实现企业RAG成本高效私有检索的案例研究
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
浏览论文内容
中文总结 AI 辅助
研究RAG系统向量检索层挑战,介绍基于TurboQuant的TurboVec开源向量索引。通过实验对比,表明TurboQuant在内存和召回率上有优势,TurboVec部署效果好,码本无关设计提升隐私性,不过存在单一数据集评估等局限。
中文摘要 AI 辅助
检索增强生成(RAG)系统为企业大语言模型应用提供了越来越多的支持,但其向量检索层存在两个未充分探索的挑战:一是训练的码本量化器在索引构建期间可能会暴露语料库统计信息,在多租户部署中形成泄漏通道;二是用于租户隔离的事后过滤会降低选择性查询的召回率。我们研究了TurboVec,这是一个基于TurboQuant构建的开源向量索引,TurboQuant是一种无需依赖语料库训练的码本无关标量量化器。在DBpedia OpenAI嵌入基准测试(d=1536,100K - 999K向量)中,TurboQuant 4位在相同内存预算下,在所有规模上的Recall@5比训练的FAISS乘积量化高出8.5 - 8.9个百分点。与HNSW(R@5 = 0.991)和IVF - PQ(R@5 = 0.840)相比,TurboQuant占据了一个独特的设计点:在不进行训练的情况下比IVF - PQ具有更高的召回率,内存比HNSW少4 - 8倍。部署在Snowpark容器服务上,TurboVec在100K向量时的中位数查询延迟为11毫秒,而仓库暴力扫描为707毫秒。内核级允许列表过滤在10 - 1000个租户工作负载中保持0.86 - 0.93的Recall@10,而后过滤基线为0.09 - 0.19。码本无关设计将成员推理准确率降低到接近随机水平(50.0%),而PQ码本为57.3%。局限性包括单一数据集评估、未压缩HNSW比较以及仅对合成数据进行隐私评估。
英文摘要
Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two underexplored challenges: (1) trained codebook quantizers may expose corpus statistics during index construction, creating a leakage channel in multi-tenant deployments, and (2) post-hoc filtering for tenant isolation degrades recall on selective queries. We study TurboVec, an open-source vector index built on TurboQuant - a codebook-oblivious scalar quantizer requiring no corpus-dependent training. On the DBpedia OpenAI embeddings benchmark (d=1536, 100K-999K vectors), TurboQuant 4-bit outperforms trained FAISS Product Quantization at the same memory budget by 8.5-8.9 percentage points in Recall@5 across all scales. Compared to HNSW (R@5=0.991) and IVF-PQ (R@5=0.840), TurboQuant occupies a distinct design point: higher recall than IVF-PQ without training, at 4-8x less memory than HNSW. Deployed on Snowpark Container Services, TurboVec achieves 11ms median query latency at 100K vectors versus 707ms for warehouse brute-force scan. Kernel-level allowlist filtering maintains 0.86-0.93 Recall@10 across 10-1000 tenant workloads versus 0.09-0.19 for post-filter baselines. Codebook-oblivious design reduces membership inference accuracy to near-random (50.0%) versus 57.3% for PQ codebooks. Limitations include single dataset evaluation, uncompressed HNSW comparison, and privacy evaluation on synthetic data only.