AI 中文总结
本文提出OpRAG,一种面向GPU多阶段RAG的资源确定性分布式运行时,通过优化编排层降低非模型开销,在多模型多基准测试中显著提升RAG端到端性能。
AI 中文摘要
智能体检索增强生成(RAG)系统结合了预处理、嵌入、检索、内存访问、上下文构建、生成和向量索引更新等环节。尽管大语言模型(LLM)解码受限于GPU,但围绕其的编排层仍可能因序列化开销、碎片化调度、低效批处理以及CPU-GPU流水线停顿而限制端到端性能。现有框架提供灵活的控制流,分布式运行时则提供可扩展的任务并行性,但二者均未将RAG阶段暴露为具有确定性执行语义的资源感知算子。本文提出OpRAG,这是一种面向GPU支持的多阶段RAG工作流的资源确定性分布式运行时。OpRAG将嵌入、检索、推理、内存和更新操作建模为一等算子,并将其降级为感知通信的执行图。它结合了Arrow零拷贝数据平面、持久工作进程、有界队列、CPU分词器预取、批处理GPU嵌入以及重叠检索/生成执行,以减少LLM推理周围的非模型开销。我们使用Llama3-8B和Mistral-7B模型,搭配FlashAttention~2、BF16执行以及32K RAG块进行评估。在端到端GPU流水线实验中,OpRAG相比最接近的竞争对手,Llama3-8B提升16.16%,Mistral-7B提升15.66%;相比RayScalableRAG,分别提升20.57%和20.71%。相较于LangChain、LangGraph、CrewAI和AutoGen,OpRAG比最佳框架基线快17.77%和17.48%。在Higress风格的查询服务中,OpRAG将混合检索延迟降低59.20%至59.62%,生成场景延迟降低52.48%至53.55%,同时保持100%的Recall@5。这些结果表明,优化分布式编排层可在不修改LLM解码内核的情况下,显著提升GPU支持的多阶段RAG性能。
英文摘要
Agentic retrieval-augmented generation (RAG) systems combine preprocessing, embedding, retrieval, memory access, context construction, generation, and vector-index updates. Although LLM decoding is GPU-bound, the surrounding orchestration layer can still limit end-to-end performance through serialization overhead, fragmented scheduling, inefficient batching, and CPU--GPU pipeline stalls. Existing frameworks provide flexible control flow, while distributed runtimes provide scalable task parallelism, but neither exposes RAG stages as resource-aware operators with deterministic execution semantics. We present OpRAG, a resource-deterministic distributed runtime for GPU-backed multi-stage RAG workflows. OpRAG models embedding, retrieval, reasoning, memory, and upsert as first-class operators and lowers them into communication-aware execution graphs. It combines an Arrow zero-copy data plane, persistent workers, bounded queues, CPU tokenizer prefetching, batched GPU embedding, and overlapped retrieval/generation execution to reduce non-model overhead around LLM inference. We evaluate OpRAG using Llama3-8B and Mistral-7B with FlashAttention~2, BF16 execution, and 32K RAG chunks. In end-to-end GPU pipeline experiments, OpRAG improves over the nearest competitor by 16.16% for Llama3-8B and 15.66% for Mistral-7B, and over RayScalableRAG by 20.57% and 20.71%, respectively. Against LangChain, LangGraph, CrewAI, and AutoGen, OpRAG is 17.77% and 17.48% faster than the best framework baseline. In Higress-style query serving, OpRAG reduces hybrid retrieval latency by 59.20--59.62% and generation-scenario latency by 52.48--53.55%, while preserving 100% Recall@5. These results show that optimizing the distributed orchestration layer can substantially improve GPU-backed multi-stage RAG without modifying the LLM decoding kernel.
Comments14 pages, 6 figures, 5 Tables