发表机构
University of California, San Diego(加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对RAG密集向量检索性能瓶颈,提出软硬件协同设计的存储内检索加速器D-NOVA。通过深度嵌入搜索功能、采用新距离度量及引入适配器,实现高效检索,相比CPU和现有加速器有显著性能和能效提升,展现全存储向量搜索加速RAG的潜力。
AI 中文摘要
检索增强生成(RAG)通过从外部知识库中检索相关信息来增强大语言模型(LLM)推理的事实基础。但其密集向量检索带来显著延迟和能量开销,成为主要性能瓶颈。近期存储内加速器虽旨在减少数据移动,但仍依赖内存外部的主机或嵌入式处理器,无法完全克服带宽限制。为此提出D-NOVA,一种软硬件协同设计的存储内检索加速器。它通过将搜索功能深度嵌入NAND内存阵列执行基于倒排文件(IVF)的分层检索管道,采用新距离度量双边界紧密相似性感知(DTS),还引入轻量级对比适配器。实验表明,D-NOVA比CPU基线快41.7倍、节能71倍,比现有存储内RAG加速器吞吐量高12.13倍、节能1.26倍,展现了全存储向量搜索对可扩展RAG加速的潜力。
英文摘要
Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces significant latency and energy overhead, becoming the primary performance bottleneck. Although recent in-storage accelerators aim to reduce data movement, they still rely on host or embedded processors outside the memory, where nearly 70% of the total retrieval time is spent. As a result, they cannot fully overcome the bandwidth limitations, leading to yet another memory bottleneck. To tackle these limitations, we present D-NOVA, a hardware-software co-designed in-storage retrieval accelerator. D-NOVA executes an inverted file (IVF)-based hierarchical retrieval pipeline by deeply embedding the search functionality directly into the NAND memory array. This is achieved by incorporating a new distance metric, Dual-Bound Tight Similarity Sensing (DTS), which is specifically tailored for searching within the NAND string. In addition, we introduce a lightweight contrastive adapter that maps embedding vectors into a DTS-friendly domain, recovering near-software recall while improving performance and energy efficiency. D-NOVA is up to 41.7x faster and 71x more energy-efficient than a CPU baseline, and achieves 12.13x higher throughput while being up to 1.26x more energy-efficient than state-of-the-art in-storage RAG accelerators, demonstrating the potential of fully in-storage vector search for scalable RAG acceleration.
CommentsAccepted at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026), Athens, Greece. Chang Eun Song and Sumukh Pinge are co-first authors and contributed equally