PReM:学习保留内容及何时刷新以进行上下文压缩
PReM: Learning What to Preserve and When to Refresh for Context Compression
浏览论文内容
中文总结 AI 辅助
研究如何进行上下文压缩,提出PReM框架,将长上下文作为模型内部KV内存,用专用内存层决策、特殊令牌触发刷新,引入相分离刷新训练,实验表明该框架在压缩下性能优,平衡了答案质量和推理效率。
中文摘要 AI 辅助
高效的长上下文推理不仅关乎降低内存成本,还在于随着生成过程保持有用的上下文证据可访问。然而,现有的面向压缩的方法,如键值(KV)缓存压缩和上下文压缩,往往要么过早决定保留哪些上下文信息,要么依赖外部压缩器。这使得难以使压缩后的上下文适应后续推理步骤所需的证据。本文介绍了PReM(保留并刷新内存),一个上下文压缩框架,它将长上下文作为模型的内部逐层KV内存来维护,并学习保留什么以及何时刷新它。具体而言,PReM使用专用内存层进行内存选择决策,并使用特殊内存令牌<m>在生成期间触发刷新。为训练这种行为,PReM引入了相分离刷新训练,在保留刷新间连续性的同时,使内存选择与内存条件生成对齐。对32K令牌上下文的实验表明,PReM在16倍和32倍压缩下均优于强大的基线,同时在答案质量和推理效率之间保持良好平衡。
英文摘要
Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds. However, existing compression-oriented approaches, such as key-value (KV) cache compression and context compression, often either make an early decision about which contextual information to keep or rely on an external compressor. Such designs make it difficult to adapt the compressed context to the evidence needed by later reasoning steps. This paper introduces PReM (Preserve and Refresh Memory), a context-compression framework that maintains the long context as the model's internal layer-wise KV memory and learns what to preserve and when to refresh it. Specifically, PReM uses a dedicated memory layer to make memory-selection decisions, and a special memory token <m> to trigger refreshes during generation. To train this behavior, PReM introduces Phase-Separated Refresh Training, aligning memory selection with memory-conditioned generation while preserving continuity across refreshes. Experiments with 32K-token contexts show that PReM outperforms strong baselines under both 16x and 32x compression, while maintaining a favorable balance between answer quality and inference efficiency.