arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23553cs.DB

Chimera:基于GPU-CPU协同处理的高效多向量检索

Chimera: Efficient Multi-Vector Retrieval via GPU-CPU Co-Processing

Yanqi Chen, Juelin Liu, Alexandra Meliou, Xiao Yan

首次发表
浏览论文内容

中文总结 AI 辅助

Chimera是一种GPU-CPU协同处理的多向量检索系统,通过在GPU存量化码、CPU存高精度数据消除数据传输瓶颈,在相同召回率下QPS较现有方法最高提升16.0倍。

中文摘要 AI 辅助

多向量检索已成为需要细粒度匹配的语义搜索应用的核心基础组件,例如信息检索、推荐系统和生物信息学领域。然而,其较高的计算复杂度和内存开销使得低延迟检索难以实现。现有系统已尝试优化查询延迟,但其设计仍以CPU为中心。尽管GPU具备显著的计算优势,但其有限的内存容量要求采用异构架构,即数据集存储在主机内存中,GPU作为加速器。现有基于GPU的系统PLID存在CPU-GPU数据传输瓶颈,因为向量数据必须在查询时从主机内存传输到GPU。我们提出Chimera,一种用于多向量检索的GPU-CPU协同处理系统,消除了该传输瓶颈。Chimera将高度压缩的低精度量化码存储在GPU上,同时在CPU内存中维护高精度数据。在查询时,它利用驻留在GPU中的数据进行高效的候选生成与过滤,并通过GPU-CPU协同评分方案进一步优化结果,该方案完全避免了向量数据传输,同时实现了计算重叠。在真实世界数据集上的实验表明,Chimera显著优于现有方法,在相同召回率水平下实现了高达16.0倍的QPS提升。

英文摘要

Multi-vector retrieval has become an important primitive for fine-grained matching in information retrieval, with emerging applications in areas such as recommender systems and bioinformatics. However, its high computational complexity and memory costs make low-latency retrieval difficult. Prior systems have attempted to optimize query latency, but their designs remain CPU-centric. While GPUs offer substantial computational advantages, their limited memory capacity necessitates a heterogeneous architecture in which the dataset resides in host memory and the GPU serves as an accelerator. Existing GPU-based system, PLAID, is bottlenecked by CPU-GPU data movement, as vector data must be transferred from host memory to the GPU at query time. We propose Chimera, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck. Chimera stores highly compressed, low-precision quantization codes on the GPU while maintaining high-precision data in CPU memory. At query time, it leverages GPU-resident data for efficient candidate generation and filtering, and further refines results through a GPU-CPU collaborative scoring scheme that completely avoids vector data transfer while enabling computation overlap. Experiments on real-world datasets demonstrate that Chimera significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.

↑