arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编解码器规范:为Transformer KV缓存学习压缩友好型规范

Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches

Yitao Jiang, Yaoqing Yang, Luyang Zhao, Muhao Chen, Devin Balkcom

arXiv 2607.20538首次发表:更新:

AI 中文总结

研究针对Transformer KV缓存压缩问题,提出Codec-Gauge方法,通过训练后缓存坐标层学习正交通道变换,结合频率分布目标优化。实验表明该方法能有效降低散度、提升量化质量,确立缓存坐标几何为提高压缩保真度的实用训练后变量。

AI 中文摘要

长上下文Transformer推理越来越依赖于KV缓存压缩或量化。先前的旋转和变换编码结果表明,每个键/值向量的通道基础会影响固定后端保留模型行为的忠实程度。我们引入了Codec-Gauge,这是一个训练后缓存坐标层,它围绕现有的压缩和量化后端学习小的正交通道变换。其频率分布目标将令牌通道DCT频谱质心损失与平滑速率代理相结合,以将KV能量集中在面向编解码器的低频布局中。我们使用测量字节和滚动压缩历史评分来评估实际压缩和解压缩。在6个模型中,3、4和6位/值的学习规范相对于原始坐标平均将zfp KL散度降低了44.0%,并且优于随机、哈达玛、DCT和PCA/KLT控制。相同的规范提高了块均匀和KIVI风格量化的质量保留。在27B模型和长上下文任务提示上的实验重现了质量趋势。这些结果将缓存坐标几何确立为一个实用的训练后变量,用于在不改变模型权重、注意力语义或后端编码规则的情况下提高压缩保真度。

英文摘要

Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results suggest that the channel basis of each key/value vector affects how faithfully a fixed backend preserves model behavior. We introduce Codec-Gauge, a post-training cache-coordinate layer that learns small orthogonal channel transforms around existing compression and quantization backends. Its frequency-distribution objective combines a token-channel DCT spectral-centroid loss with a smooth rate proxy to concentrate KV energy in low-frequency codec-facing layouts. We evaluate actual compression and decompression using measured bytes and rolling compressed-history scoring. Across six models at $3$, $4$, and $6$ bits/value, learned gauges reduce zfp KL divergence by $44.0\%$ on average relative to raw coordinates and outperform random, Hadamard, DCT, and PCA/KLT controls. The same gauges improve quality preservation for block-uniform and KIVI-style quantization. Experiments on a 27B model and long-context task prompts reproduce the quality trend, while serial storage and timing measurements validate the implemented compressed-cache paths. These results establish cache-coordinate geometry as a practical post-training variable for improving compression fidelity without changing model weights, attention semantics, or backend coding rules.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑