arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExaServe:百亿亿次高性能计算系统上的大规模LLM服务

ExaServe: Large-Scale LLM Serving on Exascale HPC Systems

Wenyi Wang, Shu Shi, Yadu Babuji, Ian Foster, Kyle Chard

arXiv 2609.10812首次发表:更新:

AI 中文总结

ExaServe是一个pip可安装的框架,将YAML规范转换为可复现的大规模LLM服务部署,在Aurora上从1到256节点实现近线性扩展,达到27.1k请求/秒,并识别了Ray Serve的O(N²)控制平面瓶颈。

AI 中文摘要

云原生的LLM服务框架已使数据中心中的部署成为常规操作,但在顶级超级计算机上部署它们仍然是一项工程挑战,需要调度器集成、MPI启动、加速器选择、节点本地权重暂存以及平台特定补丁。我们提出了ExaServe,一个可通过pip安装的框架,它将声明式YAML规范转换为可复现的大规模LLM服务部署。使用ExaServe,我们在ALCF Aurora上从1个节点扩展到256个节点(3072个vLLM副本)部署了LLM服务。非流式推理几乎线性扩展到256个节点,达到每秒27.1千个请求(每秒380万个令牌)。令牌流式传输的扩展方式不同:集中式代理在每秒约4.7千个请求时达到平台期,尽管模型服务器仍保持在服务级别目标内。我们还发现了一个O(N²)的Ray Serve控制平面瓶颈,该瓶颈将集群启动时间增加到256个节点时约30分钟。ExaServe提供了一条实用、可复现的部署路径,同时揭示了未来百亿亿次LLM服务的关键障碍。

英文摘要

Cloud-native LLM serving frameworks have made deployment routine in data centers, yet deploying them on leadership-class supercomputers remains an engineering challenge requiring scheduler integration, MPI launch, accelerator selection, node-local weight staging, and platform-specific patches. We present ExaServe, a pip-installable framework that transforms a declarative YAML specification into a reproducible large-scale LLM serving deployment. Using ExaServe, we deploy LLM serving on ALCF Aurora from 1 to 256 nodes (3072 vLLM replicas). Non-streaming inference scales nearly linearly to 256 nodes, reaching 27.1k requests/s (3.8M tokens/s). Token streaming scales differently: a centralized proxy plateaus at about 4.7k requests/s despite the model servers remaining within the service-level objective. We also identify an O(N^2) Ray Serve control-plane bottleneck that increases cluster bring-up to roughly 30 minutes at 256 nodes. ExaServe provides a practical, reproducible deployment path while exposing key barriers to future exascale LLM serving.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑