arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29913cs.CL

MILO:基于块级低秩压缩的高效多示例上下文学习

MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

  • University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

Youpeng Zhao, Tian Tan, Liqian Peng, Jun Wang, Alec Go

AI总结:

MILO提出块级低秩压缩方法,动态分配秩预算压缩KV缓存,在Qwen2.5上实现50%内存减少和1.8倍吞吐提升,性能几乎无损。

AI中文摘要:

多示例上下文学习(Many-shot in-context learning, ICL)使大型语言模型(LLMs)能够通过依赖数千个演示示例来适应复杂任务,但这种范式将推理效率的瓶颈转移到了键值(KV)缓存内存上。由于KV缓存具有线性扩展行为,存储这些中间张量已成为在线服务和设备端部署面临的首要挑战。为解决此问题,我们提出了一种新颖的压缩框架,称为MILO,该框架利用多示例上下文中固有的低秩冗余。具体而言,MILO采用块级低秩压缩策略,在块粒度上压缩KV缓存,其中每个块包含多个多示例示例。此外,为处理不同块之间异构的上下文密度,MILO基于信息熵动态分配秩预算,在积极压缩冗余块的同时保留关键块的保真度。在Qwen2.5模型上的实验结果表明,我们的方法在分类和推理基准上实现了高达50%的KV缓存内存减少和1.8倍的吞吐量提升,且性能下降可忽略不计,显著优于先前基线。

英文摘要:

Many-shot in-context learning (ICL) enables large language models (LLMs) to adapt to complex tasks by conditioning on thousands of demonstration examples, but this paradigm shifts the inference efficiency bottleneck to the key-value (KV) cache memory. Due to the linear scaling behavior of the KV cache, storing these intermediate tensors has become a paramount challenge for both online serving and on-device deployment. To address this issue, we propose a novel compression framework, termed MILO, that exploits the low-rank redundancy inherent in many-shot contexts. Specifically, MILO features a block-wise low-rank compression strategy that compresses the KV cache at the block granularity, where each block contains multiple many-shot examples. Furthermore, to handle the heterogeneous context density across different blocks, MILO dynamically allocates rank budgets based on the information entropy, preserving the fidelity of critical blocks while aggressively compressing redundant ones. Experimental results on Qwen2.5 models demonstrate that our method achieves up to 50% reduction in KV cache memory and 1.8x throughput improvement, with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.

补充信息

↑