DASC:面向混合线性注意力服务的感知衰减状态压缩
DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
浏览论文内容
中文总结 AI 辅助
该研究针对混合线性注意力服务的缓存管理问题,提出DASC方法,通过感知衰减压缩状态检查点,在保持接近全缓存质量的同时,降低内存占用并提升推理效率。
中文摘要 AI 辅助
混合线性注意力架构近期已扩展至大开放权重模型,在提供与全注意力相当质量的同时,大幅降低了键/值(KV)缓存的增长。然而,其原位循环状态更新使缓存管理复杂化:前缀复用需要状态检查点与全注意力KV,而存储状态检查点会增加内存压力,导致更多驱逐和重复预填充。通过分析门控DeltaNet(GDN)和Kimi Delta注意力(KDA)的衰减结构,我们发现不同头和通道保留前缀信息的时间尺度显著不同,我们将此称为“保留 horizon”。这种差异表明持久状态检查点存在巨大压缩潜力。基于此观察,我们提出“感知衰减状态压缩(DASC)”,该方法从模型权重中推导保留horizon,选择长horizon状态单元,并将其打包为不规则状态检查点布局。为与张量并行推理引擎高效集成,DASC进一步在张量并行(TP)秩间平衡压缩后的状态检查点。在复用阶段,DASC要么用零填充省略的单元,要么以额外计算成本从有界后缀刷新这些单元。在Kimi-Linear的检索和端到端推理基准测试中,保守的DASC配置接近全缓存,同时将KDA循环状态检查点压缩2.63倍。在固定状态检查点内存预算下,由此产生的容量增益将平均首 token 时间(TTFT)降低42.6%,并将输入吞吐量提高68.4%。在更大压缩比下,后缀刷新可恢复因更激进省略而损失的大部分准确性,代价是额外的重放计算。带有GDN的Qwen表现出类似的质量-效率趋势,表明DASC可从通道级KDA扩展至头级GDN。
英文摘要
Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints in full increases memory pressure, leading to more evictions and repeated prefill. By analyzing the decay structure of Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), we find that different heads and channels retain prefix information over markedly different timescales, which we term \emph{retention horizons}. This variation suggests substantial compression potential in persistent state checkpoints. Building on this observation, we introduce \emph{Decay-Aware State Compression} (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout. To integrate efficiently with tensor-parallel inference engines, DASC furtherly balances compressed state checkpoints across TP ranks. On reuse, DASC either zero-fills omitted units or refreshes them from a bounded suffix with additional compute cost. Across retrieval and end-to-end reasoning benchmarks on Kimi-Linear, conservative DASC configurations remain close to full caching while compressing KDA recurrent state checkpoints by $2.63\times$. Under fixed state checkpoint memory budgets, the resulting capacity gains reduce mean Time to First Token (TTFT) by 42.6\% and improve input throughput by 68.4\%. At larger compression ratio, suffix refresh recovers much of the accuracy lost to more aggressive omission, at the cost of additional replay computation. Qwen with GDN exhibits a similar quality--efficiency trend, showing that DASC extends from channel-wise KDA to head-wise GDN.
发表机构
- East China Normal University(华东师范大学)
- Meituan(美团)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。