发表机构
University of California, Santa Barbara; Hanbat National University(加州大学圣巴巴拉分校; 韩bat国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoinRAG通过两阶段检索识别检索块内与查询相关的细粒度语义单元,复用其KV缓存优化长上下文RAG,在低延迟约束下降低运算成本并提升问答F1值。
AI 中文摘要
检索增强生成(RAG)领域近期的优化研究已利用块级KV缓存复用,避免处理长检索上下文以提升效率,但粗粒度块中仍存在大量信息冗余与噪声。本文在低预填充延迟约束下优化帕累托前沿,同时通过提出CoinRAG(面向长上下文RAG的上下文信息 nugget KV缓存复用)最大化准确率。该名称隐喻反映核心机制:如同组装小标记(或“硬币”)积累更大价值,CoinRAG以更具语义相关性且紧凑的方式,组合复用离线计算的细粒度 nugget 缓存,高效形成学习到的上下文表示。具体而言,CoinRAG不进行完整块编码,而是通过两阶段检索在检索块内识别与查询相关的语义单元,并将其切片后的KV表示与块级上下文无缝组装。在LongBench多跳问答任务上的广泛评估表明,CoinRAG显著降低了运算成本,在标准快速预填充延迟预算下,相较于其他基线实现了新的帕累托前沿,且答案质量(F1值)平均相对提升5.3%。
英文摘要
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflects our core mechanism: much like assembling small tokens (or "coins") to accumulate a larger value, CoinRAG compositionally reuses offline-computed, fine-grained nugget caches to form a learned contextual representation efficiently in a more semantically relevant but compact manner. Specifically, instead of full-chunk encoding, CoinRAG identifies query-relevant semantic units within retrieved chunks through two-stage retrieval and seamlessly assembles their sliced KV representations with a chunk-level context. Extensive evaluations on LongBench multi-hop question answering tasks demonstrate that CoinRAG significantly reduces operational costs and outperforms the other baselines with a new Pareto frontier and an average 5.3% relative improvement in answer quality (F1) under a standard fast prefill latency budget.