发表机构
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对RAG压缩的权衡问题,提出RAGOCR框架,通过查询感知动态分辨率机制压缩检索文档为视觉表征,在五个QA基准上实现比朴素RAG更高准确率和更低token需求,且优于各类压缩基线。
AI 中文摘要
检索增强生成(RAG)已成为知识密集型问答的核心技术,但扩展RAG流程面临挑战,因为处理冗长的检索上下文会产生过高的计算成本。现有压缩方法存在根本权衡:硬压缩方法以查询感知方式在线运行,但仅能实现适度压缩率,且通常需要对生成模型进行微调;软压缩方法可达到更高压缩率,但依赖成本高昂的离线编码,完全不感知输入查询。为弥合这一差距,我们提出RAGOCR,这是一种将检索文档压缩为依赖输入查询的紧凑视觉表征的新型框架。为进一步平衡压缩率与信息保真度,我们引入查询感知动态分辨率机制,该机制根据每个文档的估计相关性和复杂性自适应分配视觉粒度:高度相关的段落以更高分辨率渲染以保留细粒度细节,而外围文档则以更低分辨率进行激进压缩。在使用MedOmniKB检索语料库的五个问答基准上进行的实验表明,RAGOCR的准确率比朴素RAG高出15%以上,同时仅需要八分之一的输入token数量,并且在不同检索深度下始终优于硬压缩和软压缩基线。
英文摘要
Retrieval-Augmented Generation (RAG) has become essential for knowledge-intensive question answering, yet scaling RAG pipelines remains challenging due to the prohibitive computational cost of processing lengthy retrieved contexts. Existing compression approaches face a fundamental trade-off: hard compression methods operate online in a query-aware fashion but achieve only modest compression rates and typically require fine-tuning the generative model, while soft compression methods attain higher ratios but rely on costly offline encoding that is entirely agnostic to the input query. To bridge this gap, we introduce RAGOCR, a novel framework that compresses retrieved documents into compact visual representations conditioned on the input query. To further balance compression rate and information fidelity, we introduce a query-aware dynamic resolution mechanism that adaptively allocates visual granularity based on each document's estimated relevance and complexity: highly relevant passages are rendered at higher resolution to preserve fine-grained details, while peripheral documents are aggressively compressed at lower resolution. Experiments on five QA benchmarks using the MedOmniKB retrieval corpus demonstrate that RAGOCR surpasses naive RAG by over 15\% in accuracy while requiring only one-eighth the number of input tokens, and consistently outperforms both hard and soft compression baselines across varying retrieval depths.
CommentsUnder reviewing