arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36760cs.LGcs.CL

QuantMLA: 面向低比特 MLA KV 缓存的功能对齐双路径量化

QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong

首次发表
浏览论文内容

中文总结 AI 辅助

针对 MLA 缓存随上下文和批大小线性增长的问题,提出 QuantMLA 功能对齐双路径量化框架,实现内容与 RoPE 缓存联合 INT4 量化,最小化精度损失,并提供 3.59 倍压缩和 5.168 倍吞吐提升。

中文摘要 AI 辅助

多头潜在注意力(MLA)通过其内容和解耦的 RoPE 路径实现具有紧凑缓存的高表达力多头注意力,然而缓存内存仍随上下文长度和批处理大小线性增长。在本工作中,我们建立了 MLA 双路径量化误差的系统模型,刻画了它们对注意力输出失真的不同影响,并解释了 RoPE 路径误差的显著放大现象。在该分析的指导下,我们提出了 QuantMLA,一个用于低比特双路径量化的功能对齐框架。我们推导了路径特定的变换空间,这些空间在保持全精度计算的同时,可完全离线融合到模型参数中,从而消除了在线变换开销。在这些空间内,QuantMLA 学习具有功能对齐目标的路径特定变换:注意力输出重建捕获内容路径的耦合匹配和聚合误差,而位置 QK 重建保留注意力 logits 中 RoPE 诱导的分量,并给出了输出失真的理论界限。在四个 MLA 模型家族中,QuantMLA 实现了(据我们所知)首次报道的内容和 RoPE 缓存的联合 INT4 缓存,且精度下降最小。进一步将内容缓存压缩到 INT2,同时将 RoPE 键缓存保持在 INT4,在具有挑战性的推理和代码基准上保持了有竞争力的性能。我们开发了一个原生低比特 MLA 注意力内核,将解包和反量化直接集成到注意力计算中。物理缓存布局在 128K 上下文下提供 3.59 倍压缩,而缓存压力服务负载实现了比 BF16 高 5.168 倍的整作业输出吞吐量。代码将在接收后发布。

英文摘要

Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.

发表机构

  • The University of Hong Kong(香港大学)
  • Meituan LongCat Team(美团龙猫团队)
  • South China University of Technology(华南理工大学)
  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑