超越未压缩的压缩:RAG中软上下文压缩的两阶段训练方案
Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
浏览论文内容
中文总结 AI 辅助
针对RAG中检索上下文冗长导致推理效率低的问题,提出两阶段训练方案DEX-Comp,在5个开放域QA基准上实现16倍压缩、4-24倍加速,性能与未压缩RAG相当或更优。
中文摘要 AI 辅助
检索增强生成(Retrieval-Augmented Generation, RAG)用外部知识增强语言模型,但冗长的检索上下文会增加输入长度并降低推理效率。软上下文压缩将每个文档编码为短得多的嵌入序列。然而,现有大多数方法通过蒸馏未压缩RAG系统的输出来训练,相对于原始模型的性能存在固有局限。为解决该局限,本文提出DEX-Comp,一种两阶段训练方案:纯蒸馏(Pure Distillation)仅基于未压缩RAG的正确响应为压缩模型提供预热,硬探索(Hard Exploration)则仅对未压缩RAG出错的查询运行强化学习,迫使模型探索更适配压缩表示的计算模式。在5个开放域QA基准上,检索深度从Top-5到Top-30不等,DEX-Comp将检索上下文压缩16倍,推理加速4倍至24倍,且在各检索深度的性能与未压缩RAG基准相当或更优。跨不同数据集和 backbone 的消融实验与评估进一步验证了各阶段的作用及本方法的泛化性。
英文摘要
Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-$5$ to top-$30$, DEX-Comp at $16\times$ compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by $4.4\times$--$23.7\times$. Evaluations across additional datasets and backbones further demonstrate its generalization.
发表机构
- Shandong University(山东大学)
- Bloomberg(彭博公司)
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。