arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25537cs.CLcs.AI

将长上下文压缩为面向答案的记忆嵌入以用于LLM推理

Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM长上下文推理成本高的问题,提出CMC框架,将长上下文压缩为与冻结解码器对齐的记忆嵌入,结合两级KV缓存与答案目标蒸馏,在多个QA基准上显著提升性能并降低推理开销。

中文摘要 AI 辅助

大语言模型(LLM)推理受到自注意力二次方扩展和KV缓存线性扩展的制约,随着上下文长度的增加,推理延迟、能耗和GPU内存需求也随之上升。现有的软压缩方法要么在推理时缺乏查询引导的记忆选择,要么在没有答案目标监督的情况下进行训练,要么将压缩与特定的解码器架构紧密耦合。我们提出了一种上下文到答案对齐的记忆压缩(CMC)框架,该框架将长输入上下文压缩为紧凑的上下文记忆嵌入(CMEs),这些嵌入与任意冻结解码器的嵌入空间对齐,从而在不修改解码器权重的情况下降低推理成本。CMC引入了一个两级KV缓存,将问题引导的CME选择与局部上下文窗口相结合,并通过从冻结的LLM进行答案目标蒸馏来训练压缩器。在九种编码器-解码器组合和四个QA基准上的实验表明,CMC始终优于基线,在SQuAD上实现了高达7.3的EM和4.0的F1分数提升,同时在生成3000个token时,推理时间和能耗最多降低20%,峰值保留GPU内存最多降低50%。消融研究证实,每个架构组件和训练目标都对性能有所贡献。

英文摘要

Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection at inference time, train without answer-targeted supervision, or couple compression tightly to a specific decoder architecture. We propose a Context-to-Answer-Aligned Memory Compression (CMC) framework, which compresses long input contexts into compact Context Memory Embeddings (CMEs) aligned to any frozen decoder's embedding space, reducing inference costs without modifying decoder weights. CMC introduces a two-tier KV cache that combines question-guided CME selection with a local context window, and trains the compressor with answer-targeted distillation from a frozen LLM. Experiments across nine encoder-decoder combinations and four QA benchmarks show that CMC consistently outperforms the baseline, achieving up to 7.3 EM and 4.0 F1 point gains on SQuAD, while reducing inference time and energy consumption by up to 20% and peak reserved GPU memory by up to 50% at 3,000 generation tokens. Ablation studies confirm that each architectural component and training objective contributes to the performance.

发表机构

  • Lucy Family Institute for Data & Society, University of Notre Dame(圣母大学露西家庭数据与社会研究所)
  • The University of Aizu(会津大学)
  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

↑