发表机构
College of Engineering, Al Ain University(阿联酋大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出低秩注意力适配方法,在保持量化器固定的情况下,可恢复量化KV缓存的困惑度质量,在TinyLlama、Gemma等模型上取得不同程度的性能提升,但困惑度恢复未同步提升长上下文检索能力。
AI 中文摘要
低比特键值(KV)缓存可降低自回归解码所需的内存,但产生的质量损失取决于模型和量化器。我们保持量化器固定,将浮点缓存模型的行为蒸馏为低秩Q/K/V投影更新,同时学生模型执行物理打包的增量缓存。在三个随机种子下,4比特仿射缓存适配器在TinyLlama-1.1B上恢复了54.24%±2.47%的保留困惑度差距,在Gemma-4-12B上恢复了75.96%±4.04%。在相同的冻结NF4 Llama-3.1-8B基模型上,每个量化器各选一次验证运行,在KIVI K2V2下恢复60.42%,在KVarN K4V2下恢复37.61%,同时保留180个案例的关联检索。Gemma在官方4K/8K RULER子集上的得分,在未适配的4比特缓存时为42.80,适配后升至48.33(浮点为46.15),存在显著的任务异质性。最后,2比特秩-词元扫描将TinyLlama的2比特困惑度从576.10降至11.4000±0.0059(三个种子),而浮点为10.3988,但仅恢复180个检索案例中的11-12个。这些结果表明,低秩投影适配可在固定缓存格式下恢复保留的质量,且困惑度恢复无需恢复长上下文检索。
英文摘要
Low-bit key--value (KV) caches reduce the memory required for autoregressive decoding, but the resulting quality loss depends on the model and quantizer. We keep the quantizer fixed and distill the floating-cache model's behavior into low-rank Q/K/V projection updates while the student executes a physically packed incremental cache. Across three seeds, 4-bit affine-cache adapters recover $54.24\%\pm2.47\%$ of the held-out perplexity gap on TinyLlama-1.1B and $75.96\%\pm4.04\%$ on Gemma-4-12B. On the same frozen NF4 Llama-3.1-8B base, one validation-selected run per quantizer recovers $60.42\%$ under KIVI K2V2 and $37.61\%$ under KVarN K4V2, while preserving 180-case associative retrieval. Gemma's score on an official 4K/8K RULER subset rises from 42.80 with the unadapted 4-bit cache to 48.33 after adaptation (46.15 floating), with substantial task heterogeneity. Finally, a 2-bit rank--token sweep reduces TinyLlama's 2-bit PPL from 576.10 to $11.4000\pm0.0059$ across three seeds, versus 10.3988 floating, but restores only 11--12 of 180 retrieval cases. These results show that low-rank projection adaptation can recover held-out quality across fixed cache formats, while perplexity recovery need not restore long-context retrieval.