arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26633cs.AR

NELSSA:一种基于GPU-PNM异构系统的LLM混合长度服务框架,通过基于长度的请求放置实现

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement

Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun, Junseok Lee, Hyeongseok Gwak, Myunghyun Rhee, Euiseok Kim, Donguk Moon, Kwangsik Shin, Guseul Heo, Youn… 展开作者

Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun, Junseok Lee, Hyeongseok Gwak, Myunghyun Rhee, Euiseok Kim, Donguk Moon, Kwangsik Shin, Guseul Heo, Youngpyo Joo, Hoshik Kim, Jongse Park

首次发表
浏览论文内容

中文总结 AI 辅助

NELSSA是集成GPU与PNM加速器的LLM服务系统,通过基于长度的请求放置处理混合长度工作负载,解码吞吐量最高提升5.5倍,P99延迟最高降低15倍,为LLM基础设施提供新范式。

中文摘要 AI 辅助

现代大语言模型(LLM)及其智能体应用拓宽了服务工作负载的范围,上下文长度从数百个token到数十万个token不等。由于这些请求在同一服务窗口内频繁交错,LLM服务系统必须处理高度异构的混合长度工作负载。此类混合长度工作负载暴露了以GPU为中心的服务架构的基本低效性,其吞吐量依赖于受内存限制的大批次处理。本文提出NELSSA,一种将GPU与真实近内存处理(PNM)加速器设备集成的LLM服务系统,以高效支持混合长度工作负载。NELSSA采用基于长度的请求放置策略,将短上下文请求路由至GPU,长上下文请求路由至PNM层,同时纳入运行时迁移机制,以适应动态上下文增长而无需重新计算。我们将NELSSA原型化为端到端系统,在支持RPC和RDMA的CXL基础设施上实现了PNM设备级稀疏注意力、GPU解码内核,以及用于调度和跨层内存移动的主机端运行时。在混合长度LLM工作负载中,与仅使用GPU的基线系统相比,NELSSA的解码吞吐量提升最高达5.5 token/sec,P99延迟降低最高达15倍。我们的端到端原型和实验结果表明,由基于CXL的解聚技术实现的集成GPU-PNM服务,是支持不断发展的工作负载的可扩展、灵活LLM基础设施的有前景系统范式。

英文摘要

Modern LLMs and their agentic applications are broadening the range of serving workloads, spanning context lengths from a few hundred tokens to hundreds of thousands. As these requests frequently interleave within the same serving window, LLM serving systems must handle highly heterogeneous mixed-length workloads. Such mixed-length workloads expose fundamental inefficiencies in GPU-centric serving architectures, whose throughput depends on large, memory-constrained batches. In this paper, we present NELSSA, an LLM serving system that integrates GPUs with real-world Processing-near-Memory (PNM) accelerator devices to efficiently support mixed-length workloads. NELSSA employs length-based request placement to route short-context requests to GPUs and long-context requests to the PNM tier, incorporating runtime migration to accommodate dynamic context growth without recomputation. We prototype NELSSA as an end-to-end system, implementing device-level sparse attention on PNM, GPU decode kernels, and a host-side runtime that orchestrates scheduling and cross-tier memory movement over a CXL-enabled infrastructure with RPC and RDMA support. Across mixed-length LLM workloads, NELSSA improves decode throughput by up to 5.5x in tokens/sec and reduces P99 latency by up to 15x compared to GPU-only baselines. Our end-to-end prototype and experimental results suggest that integrated GPU-PNM serving, enabled by CXL-based disaggregation, is a promising system paradigm for scalable and flexible LLM infrastructures that support evolving workloads.

补充信息

↑