RAGMark:面向检索增强生成系统的综合基准测试框架
RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems
浏览论文内容
中文总结 AI 辅助
提出模块化基准框架RAGMark,用于小规模多GPU环境下评估RAG系统各组件,揭示上下文缩减与重排序的复合效益及跨阶段交互,能耗可降66%。
中文摘要 AI 辅助
我们提出了RAGMark,一个面向先进检索增强生成(RAG)系统的模块化基准测试框架,其目标环境为小规模多GPU环境。RAGMark评估多种RAG组件,包括检索器、向量数据库、提示处理方法以及生成模型,同时收集详细的逐阶段指标,如延迟、GPU利用率、内存消耗、功耗、首令牌时间(TTFT)、吞吐量和答案质量。该框架高度可扩展,将RAG阶段、计时和资源监控分离为模块化组件,并设计用于高效扫描大型配置空间,同时最小化重复的模型和数据库初始化开销。利用RAGMark,我们在开放域问答数据集上对五种RAG工作负载进行了特征化分析,涵盖了不同的检索深度、模型规模、重排序、压缩方法和向量数据库配置。我们表明,虽然自回归生成在朴素流水线中主导延迟,但上下文缩减技术将瓶颈转移到计算、内存带宽和预处理阶段。重排序和压缩产生复合效益:重排序本身减少了压缩工作负载,而两者共同降低了预填充和KV缓存遍历成本,使能耗降低高达66%。我们进一步观察到强烈的跨阶段交互,其中上游上下文的微小缩减会级联影响下游的延迟、内存流量和能耗。RAGMark源代码公开于:此https URL。
英文摘要
We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments. RAGMark evaluates diverse RAG components, including retrievers, vector databases, prompt-processing methods, and generator models, while collecting detailed per-stage metrics such as latency, GPU utilization, memory consumption, power usage, time to first token (TTFT), throughput, and answer quality. The framework is highly extensible, separating RAG stages, timing, and resource monitoring into modular components, and is designed to efficiently sweep large configuration spaces while minimizing repeated model and database initialization overhead. Using RAGMark, we characterize five RAG workloads on open-domain QA datasets across varying retrieval depths, model scales, reranking, compression methods, and vector database configurations. We show that while autoregressive generation dominates latency in naive pipelines, context-reduction techniques shift bottlenecks across compute, memory bandwidth, and preprocessing stages. Reranking and compression produce compounding benefits: reranking reduces compression workload itself, while both jointly reduce prefill and KV-cache traversal costs, lowering energy consumption by up to 66%. We further observe strong cross-stage interactions, where small upstream context reductions cascade through downstream latency, memory traffic, and energy consumption. The RAGMark source code is publicly available at: https://github.com/zferic/RAGMark.
发表机构
- Northeastern University(东北大学)
- College of William & Mary(威廉与玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。