发表机构
Peking University; Alibaba Group(北京大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有上下文并行机制的静态配置问题,提出自适应上下文并行系统Vertumnus,通过请求级路由与集群级节点调整结合全局缓存管理,在64-GPU集群上实现了平均TTFT降低28.1%、SLO达成率提升13.3个百分点的性能。
AI 中文摘要
随着大语言模型(LLM)上下文窗口扩大和输入序列变长,服务系统面临日益增长的计算与内存需求。上下文并行(CP)通过在多个计算节点(rank)间划分输入序列以并行化计算,因此对高效的LLM服务愈发重要。然而,现有支持CP的系统要么依赖静态CP配置,要么仅针对活跃请求或批次调整CP度。本文提出Vertumnus,一种为异构且动态变化的工作负载设计的自适应CP服务系统。在请求级,Vertumnus利用包含预测排队延迟、感知缓存的预填充时间及GPU时间成本的放置代价,在具有不同CP度的工作节点间路由请求;在集群级,Vertumnus通过秒级的拆分与合并操作,根据工作负载需求调整工作节点组成。Vertumnus还引入全局前缀缓存管理策略,在CP度相同或不同的工作节点间协调缓存放置与复制,在请求分配和工作节点组成变化时保持缓存局部性。在64-GPU集群上使用公开及生产工作负载开展的实验表明,在最高评估负载下,Vertumnus较最强基线将平均首字符生成时间(TTFT)降低最多28.1%,并将令牌加权的服务水平目标(SLO)达成率提升最多13.3个百分点。
英文摘要
As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the input sequence across multiple ranks to parallelize the computation, has therefore become increasingly important for efficient LLM serving. However, existing CP-enabled systems either rely on static CP configurations or adjust the CP degree only for active requests or batches. In this paper, we present Vertumnus, an adaptive CP serving system designed for heterogeneous and evolving workloads. At the request level, Vertumnus routes requests among workers with different CP degrees using a placement cost that combines predicted queuing delay, cache-aware prefill time, and GPU-time cost. At the cluster level, Vertumnus adapts the worker composition through seconds-scale split and merge operations as workload demand changes. Vertumnus further introduces a global prefix-cache management policy that coordinates cache placement and replication among workers with the same or different CP degrees, preserving cache locality as request assignments and worker composition change. Experiments on a 64-GPU cluster with public and production workloads show that, under the highest evaluated loads, Vertumnus reduces mean TTFT by up to 28.1% and improves token-weighted SLO attainment by up to 13.3 percentage points over the strongest baseline.