arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23176cs.IR

评估现代RAG:文本、多模态、密集型与后期交互管道

Evaluating Modern RAG: Textual, Multimodal, Dense, and Late Interaction Pipelines

Emre Kuru, Mehmet Onur Keskin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对传统RAG在含视觉元素文档上的局限,提出数据驱动的RAG管道选择方法,评估了文本、多模态等现代RAG管道的性能与效率权衡,为实际应用提供指导。

中文摘要 AI 辅助

检索增强生成(RAG)系统传统上依赖基于文本的管道,从文档中提取和检索信息。这类方法虽高效且轻量,但在含义通过布局、表格和视觉元素传达的文档上常表现不佳。近期由视觉语言模型(VLMs)驱动的多模态管道进展,通过联合编码视觉与文本信号提升了检索质量,但计算和内存成本有所增加。我们提出一种定量、数据驱动的选择方法,指导从业者根据给定文档语料库的经验效果与资源约束选择最合适的RAG管道。我们评估了当代文本和多模态管道,包括密集型与后期交互架构,分析了它们的权衡,并为平衡检索性能与系统效率提供可操作的指导。

英文摘要

Retrieval-augmented generation (RAG) systems have traditionally relied on text-based pipelines that extract and retrieve information from documents. While efficient and lightweight, these approaches often struggle with documents where meaning is conveyed through layout, tables, and visual elements. Recent advances in multimodal pipelines, powered by vision-language models (VLMs), improve retrieval quality by jointly encoding visual and textual signals, but at increased computational and memory cost. We propose a quantitative, data-driven selection methodology that guides practitioners in choosing the most appropriate RAG pipeline for a given document corpus based on empirical effectiveness and resource constraints. We evaluate contemporary textual and multimodal pipelines, including dense and late-interaction architectures, analyze their trade-offs, and provide actionable guidance for balancing retrieval performance with system efficiency.

↑