AI 中文总结
AdaMem提出相关性引导的软压缩框架,按查询相关性动态分配记忆令牌预算,在开放域QA基准上优于均匀分配基线,显著提升压缩效率与答案质量。
AI 中文摘要
检索增强生成(RAG)通过检索到的证据改进语言模型,但处理大量长段落成本高昂,且可能引入干扰信息。软压缩通过在生成前将段落编码为连续记忆嵌入的紧凑序列来应对这一挑战。然而,现有方法通常为每个保留的段落分配相同数量的记忆嵌入,而不考虑其与查询的相关性。为解决此问题,我们提出AdaMem,一个相关性引导的软压缩框架,将学习到的段落相关性估计映射到固定记忆令牌预算的查询相关分配上。共享的查询条件压缩器在单次传递中同时产生连续的段落记忆和相关性分数;确定性分配规则将更多记忆令牌分配给得分较高的段落,并可省略得分较低的段落。在六个开放域问答基准上,AdaMem在匹配的记忆预算下始终优于OSCAR(使用均匀分配的紧密匹配的软压缩基线)以及其他软压缩方法。在标准16倍压缩下,AdaMem相比均匀分配基线将子串匹配最多提升3.2个百分点(5.5%),平均相对增益为3.4%;在激进64倍压缩下,平均相对增益增至14.6%,在PopQA上最大提升9.8个百分点(19.7%)。AdaMem在推理延迟比全上下文基线低至4倍的情况下,匹配未压缩的答案质量。AdaMem保持了与均匀压缩基线相当的高效性,同时实现比全上下文推理低至4倍的推理延迟。因此,当检索池较大且可用记忆预算紧张时,相关性引导的记忆分配尤为有效。
英文摘要
Retrieval-augmented generation (RAG) improves language models with retrieved evidence, but processing many long passages is costly and can introduce distracting information. Soft compression addresses this challenge by encoding passages as compact sequences of continuous memory embeddings before generation. However, existing methods typically assign each retained passage an identical number of memory embeddings, irrespective of its query-specific relevance. To address this, we propose AdaMem, a relevance-guided soft-compression framework that maps learned passage-relevance estimates to a query-dependent allocation of a fixed memory-token budget. A shared query-conditioned compressor produces both continuous passage memories and relevance scores in a single pass; a deterministic allocation rule assigns more memory tokens to higher-scoring passages and can omit low-scoring ones. Across six open-domain QA benchmarks, AdaMem consistently outperforms OSCAR (the closely matched soft-compression baseline that uses uniform allocation) as well as other soft-compression methods at matched memory budgets. Under standard 16$\times$ compression, AdaMem improves sub-string match by up to 3.2 points (5.5%) over uniform allocation baseline, with an average relative gain of 3.4%; under aggressive 64$\times$ compression the average relative gain grows to 14.6%, with a maximum of 9.8 points (19.7%) on PopQA. AdaMem matches the answer quality of the uncompressed at up to 4$\times$ lower inference latency than full context baseline. AdaMem retains an efficiency profile comparable to the uniform-compression baseline, while achieving up to $4\times$ lower inference latency than full-context inference. Thus, relevance-guided memory allocation is particularly effective when retrieval pools are large and the available memory budget is tight.