块扩散语言模型中的掩码引导KV缓存驱逐
Mask-Guided KV Cache Eviction in Block Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
提出MaskAhead,一种无需训练的掩码查询排名方法,用于块扩散语言模型的KV缓存选择与驱逐,显著减少内存并加速生成。
中文摘要 AI 辅助
块扩散语言模型在生成过程中保留一个大型键值(KV)缓存,并在每个去噪步骤中对其进行关注,这限制了内存容量和生成速度。降低这些成本需要决定使用哪些过去的令牌来去噪当前块(选择)以及将哪些令牌保留在内存中以供未来块使用(驱逐)。我们提出MaskAhead,一种无需训练的方法,通过单一的掩码查询排名机制解决这两个任务。当前块的掩码指导选择,而即将到来的掩码块的探针指导驱逐。两者都根据KV条目对注意力输出的估计贡献进行排名。我们的量化变体Q-MaskAhead直接从低位KV计算选择和注意力,在很大程度上保留了所选条目。在Fast-dLLM-v2、DreamReasoner和LLaDA2.0-mini上的实验涵盖了长生成推理、长提示问答和针在海堆检索。在长提示问答中,MaskAhead平均将KV内存减少9.5倍,相对于密集推理平均F1损失为1.2点。Q-MaskAhead将减少增加到20.1倍,平均F1损失为2.3点。在批量32的系统配置文件中,MaskAhead实现了1.23倍的端到端和1.68倍的解码阶段加速,相对于密集推理。
英文摘要
Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory for future blocks (eviction). We propose MaskAhead, a training-free method that solves both tasks with a single mask-query-based ranking mechanism. Current-block masks guide selection, while probes of upcoming masked blocks guide eviction. Both rank KV entries by their estimated contribution to the attention output. Our quantized variant, Q-MaskAhead, computes selection and attention directly from low-bit KV, largely preserving the selected entries. Experiments on Fast-dLLM-v2, DreamReasoner, and LLaDA2.0-mini cover long-generation reasoning, long-prompt question answering, and needle-in-a-haystack retrieval. On long-prompt QA, MaskAhead reduces KV memory by $9.5\times$ on average with a 1.2-point mean F1 loss relative to dense inference. Q-MaskAhead increases the reduction to $20.1\times$ with a 2.3-point mean F1 loss. In a batch-32 systems profile, MaskAhead achieves $1.23\times$ end-to-end and $1.68\times$ decode-stage speedups over dense inference.