arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12121cs.CLcs.AI

QV-PIC:面向高效RAG服务的查询感知视觉位置无关缓存

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu

首次发表
浏览论文内容

中文总结 AI 辅助

QV-PIC是一种查询感知双分辨率PIC复用框架,通过离线编译视觉缓存、在线分分辨率处理,解决渲染图像PIC质量下降问题,在6项RAG任务中显著提升F1并降低TTFT。

中文摘要 AI 辅助

检索增强生成(RAG)会在不同查询间重复预填充相同的文本块,产生冗余计算。位置无关缓存(PIC)通过在不同位置间复用预计算的键值(KV)来缓解该问题,但受限于文本token数量庞大,其效率提升空间有限。将文本块渲染为图像可将文本压缩为更少的视觉token,但渲染图像的PIC相比文本PIC存在更严重的质量下降,这种特定表示的差距主要源于独立编译的缓存间的上下文不匹配,以及视觉压缩过程中细粒度文本证据的丢失。现有PIC修复方法主要通过选择性重计算解决前者,但会产生在线计算开销且无法恢复丢失的文本细节。我们提出QV-PIC,一种由模型原生模板引导的查询感知双分辨率PIC复用框架。离线阶段,QV-PIC在模型原生聊天模板前缀下编译视觉缓存,无需在线重计算即可提升PIC质量;在线阶段,它通过累积查询相关性得分,在低分辨率下保留全局上下文,并在高分辨率预算内恢复细粒度文本证据,同时保留视觉压缩的效率优势。在6项任务中,QV-PIC相比普通渲染图像PIC平均F1提升21.6个百分点,缩小了与普通文本PIC的差距,且在F1上比优化后的文本PIC高出2.58,同时将首包生成时间(TTFT)降低17.2%;相比全预填充,它将TTFT削减了83.8%。

英文摘要

Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by the large volume of text tokens. Rendering text chunks as images can compress the text into fewer visual tokens, but the rendered-image PIC suffers more severe quality degradation than the text PIC. This representation-specific gap primarily arises from contextual mismatches across independently compiled caches and the loss of fine-grained textual evidence during visual compression. Existing PIC repair methods mainly address the former through selective recomputation, but they incur online computation and cannot recover lost textual details. We propose QV-PIC, a query-aware dual-resolution PIC reuse framework guided by model-native templates. Offline, QV-PIC compiles visual caches under the model's native chat-template prefix, improving PIC quality without online recomputation. Online, it preserves global context with low resolution and restores fine-grained textual evidence within a high-resolution budget by cumulative query relevance scores, retaining the efficiency benefit of visual compression. Across six tasks, QV-PIC improves average F1 by 21.6 points over vanilla rendered-image PIC, closes the gap to vanilla text PIC, and surpasses optimized text PIC by 2.58 F1 while reducing TTFT by 17.2\%. Relative to full prefill, it cuts TTFT by 83.8%.

↑