arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeDeM:面向长上下文问答的隐藏状态记忆选择性解压

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

Maryam Haghifam, Jason Cong, Yizhou Sun

arXiv 2608.00311首次发表:更新:

发表机构

University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM长上下文推理成本高、证据利用不可靠的问题,提出SeDeM选择性解压框架,通过解耦记忆存储与解码器条件,在1B、3B主干的四个长上下文QA基准上,优于压缩基线及全上下文微调方法,还提升了解码效率。

AI 中文摘要

大型语言模型(LLM)的长上下文推理成本高昂:预填充阶段的自注意力计算复杂度随序列长度呈二次方增长,键值(KV)缓存随已处理token数量增加而膨胀。更大的上下文窗口也无法保证证据的可靠利用。上下文压缩可降低该成本,但许多软压缩方法以LLM作为压缩器,依赖紧凑的记忆token来同时保留信息并为解码器提供条件。我们提出SeDeM,一种选择性解压框架,将紧凑记忆存储与解码器条件解耦。LLM从选定的Transformer中间层提取隐藏状态,轻量压缩器将其存储为记忆块,查询条件选择器选取相关块,解压器仅将选中块扩展为与解码器中间层兼容的隐藏状态。如此,解码器既避免了全上下文处理,也无需直接从高度压缩的记忆槽生成内容。在四个长上下文问答基准上,SeDeM在1B和3B相同主干设置下均比所评估的压缩基线取得更高的问答分数,且3B主干版本在三个数据集上超越全上下文微调。学习到的选择器在训练期间使用块级证据监督。与ICAE相比,SeDeM还缩短了在线首token生成时间并提升了自回归解码吞吐量。

英文摘要

Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key-value (KV) cache grows with the number of processed tokens. Larger context windows also do not ensure reliable evidence use. Context compression reduces this cost, but many soft-compression methods use LLMs as compressors and rely on compact memory tokens both to preserve information and to condition the decoder. We propose SeDeM, a selective decompression framework that decouples compact memory storage from decoder conditioning. SeDeM stores context as compact hidden-state memory blocks, selects query-relevant blocks, and decompresses only the selected blocks for decoder conditioning. Thus, the decoder avoids both full-context processing and direct generation from highly compressed memory slots. On four long-context QA benchmarks, SeDeM achieves higher QA scores than the compression baselines in our main comparison in both 1B and 3B same-backbone settings, and with the 3B backbone exceeds full-context fine-tuning on three datasets. SeDeM also provides favorable quality--efficiency trade-offs, achieving 1.74--2.46$\times$ lower online time-to-first-token and 1.08--1.10$\times$ higher autoregressive decoding throughput relative to ICAE while maintaining strong answer quality.

CommentsAccepted at EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑