NELSSA:一种基于GPU-PNM异构系统的LLM混合长度服务框架,通过基于长度的请求放置实现
NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
浏览论文内容
中文总结 AI 辅助
NELSSA是集成GPU与PNM加速器的LLM服务系统,通过基于长度的请求放置处理混合长度工作负载,解码吞吐量最高提升5.5倍,P99延迟最高降低15倍,为LLM基础设施提供新范式。
中文摘要 AI 辅助
现代大语言模型(LLM)及其智能体应用拓宽了服务工作负载的范围,上下文长度从数百个token到数十万个token不等。由于这些请求在同一服务窗口内频繁交错,LLM服务系统必须处理高度异构的混合长度工作负载。此类混合长度工作负载暴露了以GPU为中心的服务架构的基本低效性,其吞吐量依赖于受内存限制的大批次处理。本文提出NELSSA,一种将GPU与真实近内存处理(PNM)加速器设备集成的LLM服务系统,以高效支持混合长度工作负载。NELSSA采用基于长度的请求放置策略,将短上下文请求路由至GPU,长上下文请求路由至PNM层,同时纳入运行时迁移机制,以适应动态上下文增长而无需重新计算。我们将NELSSA原型化为端到端系统,在支持RPC和RDMA的CXL基础设施上实现了PNM设备级稀疏注意力、GPU解码内核,以及用于调度和跨层内存移动的主机端运行时。在混合长度LLM工作负载中,与仅使用GPU的基线系统相比,NELSSA的解码吞吐量提升最高达5.5 token/sec,P99延迟降低最高达15倍。我们的端到端原型和实验结果表明,由基于CXL的解聚技术实现的集成GPU-PNM服务,是支持不断发展的工作负载的可扩展、灵活LLM基础设施的有前景系统范式。
英文摘要
Modern LLMs and their agentic applications are broadening the range of serving workloads, spanning context lengths from a few hundred tokens to hundreds of thousands. As these requests frequently interleave within the same serving window, LLM serving systems must handle highly heterogeneous mixed-length workloads. Such mixed-length workloads expose fundamental inefficiencies in GPU-centric serving architectures, whose throughput depends on large, memory-constrained batches. In this paper, we present NELSSA, an LLM serving system that integrates GPUs with real-world Processing-near-Memory (PNM) accelerator devices to efficiently support mixed-length workloads. NELSSA employs length-based request placement to route short-context requests to GPUs and long-context requests to the PNM tier, incorporating runtime migration to accommodate dynamic context growth without recomputation. We prototype NELSSA as an end-to-end system, implementing device-level sparse attention on PNM, GPU decode kernels, and a host-side runtime that orchestrates scheduling and cross-tier memory movement over a CXL-enabled infrastructure with RPC and RDMA support. Across mixed-length LLM workloads, NELSSA improves decode throughput by up to 5.5x in tokens/sec and reduces P99 latency by up to 15x compared to GPU-only baselines. Our end-to-end prototype and experimental results suggest that integrated GPU-PNM serving, enabled by CXL-based disaggregation, is a promising system paradigm for scalable and flexible LLM infrastructures that support evolving workloads.