发表机构
University of Illinois Urbana-Champaign; IBM Research(伊利诺伊大学厄巴纳-香槟分校; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对嵌入与生成模型孤立执行导致低吞吐和 GPU 利用率差的问题,提出统一推理循环内异构批处理的 serving 系统 Orthrus,通过分块嵌入和负载感知批处理,在 A100 上提升吞吐 1.28-4.52 倍,降低 p99 延迟达 55.8%。
AI 中文摘要
现代信息检索越来越多地同时使用嵌入模型和生成模型来处理复杂查询。然而,当前的 serving 系统由于孤立地执行这些模型,导致吞吐量低且 GPU 利用率差。粗粒度的分区(例如将 GPU 专用于特定任务)无法适应动态工作负载,并产生计算“气泡”。为了解决这些问题,我们提出了 Orthrus,一个在统一推理循环内执行异构批处理的 serving 系统。主要挑战在于统一具有冲突计算模式的嵌入和生成工作负载,同时优化批处理组合以实现高性能。Orthrus 通过分块嵌入与增量池化,并以工作负载感知的方式调整批处理组合来应对这些挑战。在四块 A100 GPU 上的评估表明,与基线部署相比,Orthrus 在受控工作负载上实现了 1.28 倍至 4.52 倍的更高吞吐量,并在迭代 RAG 基准上将端到端 p99 延迟降低了高达 55.8%。我们在 https 此 URL 发布了代码。
英文摘要
Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, fails to adapt to dynamic workloads and creates computational "bubbles". To address these, we present Orthrus, a serving system that performs heterogeneous batching within a unified inference loop. The primary challenge lies in unifying embedding and generation workloads with conflicting computational patterns while optimizing batch composition for high performance. Orthrus addresses these challenges through chunked embedding with incremental pooling and by adjusting batch composition in a workload-aware manner. Evaluation on four A100 GPUs shows that, relative to baseline deployments, Orthrus achieves 1.28$\times$--4.52$\times$ higher throughput on controlled workloads and up to 55.8% lower end-to-end p99 latency on an iterative-RAG benchmark. We release our code at https://github.com/illinoisdata/Orthrus .
Comments15 pages, 8 figures, Accepted to EMNLP 2026 (main conference)