arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21599cs.PFcs.AI

解耦注意力融合:通过高效的键值缓存重用加速检索增强生成

Decoupled Attention Fusion: Accelerating RAG with Efficient KV Cache Reuse

Xiabao Wu, Wentao Liu, Yongchao Liu, Jiajun Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

针对RAG在长上下文场景中TTFT延迟高及缓存问题,提出解耦注意力融合(DAF)框架,将注意力过程解耦为三个阶段,与Flash-Attention内核兼容,实验显示其在不牺牲准确性的情况下,相比CacheBlend和完全重新计算有显著加速。

中文摘要 AI 辅助

检索增强生成(RAG)有效减轻大语言模型中的幻觉问题,但在长上下文场景中存在较高的首次令牌时间(TTFT)延迟。重用预计算的文档键值缓存可解决此问题,但会引入分布不匹配,即离线缓存缺乏连贯推理所需的文档间注意力模式。CacheBlend通过选择性注意力减少重新计算,但在更长上下文时会严重降低准确性。为应对这些挑战,我们提出解耦注意力融合(DAF)框架,它在显著减少重新计算开销的同时保持高精度。DAF将注意力过程解耦为三个集成阶段:重要令牌自注意力恢复缺失的文档间注意力、问题-文档自注意力进行标准推理,以及状态融合连接其输出以合成最终隐藏状态。通过将这些操作解耦为密集模式,DAF与Flash-Attention内核原生兼容,无需复杂注意力掩码即可最大化硬件利用率。实验表明,在长上下文基准测试中,DAF比CacheBlend加速高达2倍,比完全重新计算加速5.6倍,且不牺牲准确性。

英文摘要

Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in Large Language Models (LLMs) but suffers from prohibitive Time-To-First-Token (TTFT) latency in long-context scenarios. Reusing pre-computed document KV caches addresses this but introduces a distribution mismatch, where offline caches lack the inter-document attention patterns required for coherent reasoning. CacheBlend reduces recomputation via selective attention, but suffers severe accuracy degradation at longer contexts. To address these challenges, we propose Decoupled Attention Fusion (DAF), a framework that maintains high accuracy while significantly reducing recomputation overhead. DAF decouples the attention process into three integrated stages: important-token self-attention to restore missing inter-document attention, question-document self-attention for standard inference, and a state fusion that concatenates their outputs to synthesize the final hidden states. By decoupling these operations into dense patterns, DAF is natively compatible with Flash-Attention kernels, maximizing hardware utilization without requiring complex attention masks. Experiments show that DAF delivers up to 2 times speedup over CacheBlend and 5.6 times over full recomputation with vLLM on long-context benchmarks, without sacrificing accuracy.

发表机构

  • Ant Group, China(蚂蚁集团,中国)
  • Southeast University, Nanjing, China(东南大学,南京,中国)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑