arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越KV重建:推测解码中MLA草稿模型的功能重建

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding

Weiye Shi, Fanxu Meng, Muhan Zhang

arXiv 2607.27269首次发表:更新:

AI 中文总结

针对推测解码中MHA/GQA转MLA导致草稿令牌接受率降低的问题,提出功能重建方法优化转换后的MLA注意力模块,在多数任务中提升了草稿令牌接受率。

AI 中文摘要

多头潜在注意力(MLA)对于长上下文大语言模型(LLM)推理愈发重要,因为紧凑的潜在状态可替代不断增长的键值(KV)缓存,减少解码内存流量。然而,多数高性能开放检查点使用多头注意力或分组查询注意力(MHA/GQA),因此需进行转换以获得MLA的缓存效率,而无需从头开始重新训练。推测解码提供互补加速,但其加速取决于草稿提议与目标验证的一致性。我们发现,直接将MHA/GQA转换为MLA会大幅降低这种一致性:低秩分解和旋转位置嵌入(RoPE)处理会引入注意力功能误差,该误差对于独立生成可能可容忍,但会显著降低草稿令牌的接受率。因此,我们将MLA草稿构建表述为功能重建,而非缓存压缩。我们的端到端(E2E)方法优化每个转换后的MLA注意力模块,使其在校隐状态上重现原始MHA/GQA对应模块经输出投影后的响应。该转换器无关的后转换过程保留转换后的缓存和推理图,既不需要验证器对数也不需要验证器监督。我们评估了192种模型-转换器-后端-方法-任务配置,涵盖4组Llama/Qwen草稿-目标对、TransMLA和MHA2MLA、HF和vLLM,以及4个200提示任务。在0.5个百分点的报告容忍度下,功能重建在64个匹配任务单元中显著提升了37个单元的接受率,使26个单元几乎不变,仅在1个单元中显著降低。代码和评估工件可在该https URL获取。

英文摘要

Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic. Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA's cache efficiency without retraining from scratch. Speculative decoding offers complementary acceleration, but its speedup depends on agreement between draft proposals and target verification. We find that direct MHA/GQA-to-MLA conversion can sharply reduce this agreement: low-rank factorization and RoPE handling introduce attention-function errors that may be tolerable for standalone generation but substantially lower draft-token acceptance. We therefore formulate MLA draft construction as functional reconstruction rather than cache compression. Our end-to-end (E2E) method optimizes each converted MLA attention module to reproduce the post-output-projection response of its original MHA/GQA counterpart on calibration hidden states. This converter-agnostic post-conversion procedure preserves the converted cache and inference graph and requires neither verifier logits nor verifier supervision. We evaluate 192 model-converter-backend-method-task configurations spanning four Llama/Qwen draft-target pairs, TransMLA and MHA2MLA, HF and vLLM, and four 200-prompt tasks. With a 0.5-percentage-point reporting tolerance, Functional Reconstruction materially improves acceptance in 37 of 64 matched task cells, leaves 26 practically unchanged, and materially decreases one. Code and evaluation artifacts are available at https://github.com/swyhahaha/FunctionalMLA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑