arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlexComp:一种模型实现上下文压缩中的任意压缩比

FlexComp: One Model for Every Ratio in Context Compression

Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa, Yoshimasa Tsuruoka

arXiv 2609.11192首次发表:更新:

发表机构

The University of Tokyo; National Institute of Informatics(东京大学; 国立情报学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FlexComp通过Matryoshka式训练和动态预算选择,实现单模型支持任意压缩比,性能接近固定比专家,显著降低KV缓存并提升吞吐量。

AI 中文摘要

软上下文压缩将上下文压缩为少量记忆令牌,由冻结的大语言模型(LLM)代替原始文本使用,但现有压缩器在训练和推理时固定压缩比:每个部署的压缩比都需要单独训练的模型,且所选压缩比统一应用于所有输入,而实际需求差异巨大。我们提出FlexComp,一种与具体方法无关的框架,将压缩比从训练和部署中解耦:Matryoshka式训练对每个实例采样记忆预算$K$,使一个模型成为任意压缩比压缩器,随后通过(1)基于置信度的级联路由或(2)轻量级学习的$K$预测器,为每个输入选择预算。在ICAE、500xCompressor和SAC上,针对MRQA数据集,单个FlexComp模型与分别训练的固定压缩比专家模型匹配,性能退化极小。级联路由在平均压缩比高达266倍时,保留了最温和压缩比准确率的98%以上;$K$预测器在单次压缩-解码过程中,达到158-236倍压缩比,与最温和压缩比的F1差距在0.7以内。在服务规模批处理大小下,$K$预测器将上下文KV缓存减少50%,并将解码吞吐量提高47%。

英文摘要

Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose actual needs vary drastically. We propose FlexComp, a method-agnostic framework that decouples the ratio from both training and deployment: Matryoshka-style training samples the memory budget $K$ per instance, turning one model into an any-ratio compressor, and the budget is then chosen per input by: (1) confidence-based cascade routing or (2) a lightweight learned $K$ predictor. Across ICAE, 500xCompressor, and SAC on MRQA, a single FlexComp model matches separately trained fixed-ratio specialists with minimal degradation. Cascade routing preserves over 98% of the mildest ratio's accuracy at up to 266x average compression; the $K$ predictor, in a single compression-decoding pass, reaches 158-236x within 0.7 F1 of the mildest ratio. At serving-scale batch sizes, the $K$ predictor cuts context KV cache by 50% and improves decoding throughput by 47%.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑