发表机构
McGill University; MBZUAI; Nutanix, Inc.(麦吉尔大学; 穆罕默德·本·扎耶德人工智能大学; Nutanix公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeltaServe是一种主机无关式LLM推理与微调协同服务设计,可利用闲置GPU算力提升LoRA微调吞吐量,已集成至vLLM等框架,在保持100%推理SLO合规的前提下,微调吞吐量较现有方案显著提升。
AI 中文摘要
大语言模型(LLM)服务系统为满足严格的延迟目标会按峰值负载配置资源,当流量低于峰值时,会留下大量闲置的GPU算力。本文提出DeltaServe,一种主机无关式协同服务设计,可在保留推理服务水平目标(SLO)的同时,将闲置的推理算力转化为LoRA微调吞吐量。DeltaServe通过仅需多LoRA批处理支持的紧凑钩子接口与现有推理引擎集成,利用推理预填充和LoRA微调前向传播的共享执行结构,并采用感知SLO的调度器,仅在有足够推理余量时才允许和执行微调。该调度器由离线校准、在线优化的感知CUDA图的延迟模型驱动。我们将DeltaServe与vLLM、SGLang和S-LoRA集成,基于某公司X的生产流量轨迹,在100%推理SLO合规的情况下,vLLM上的DeltaServe微调吞吐量是LLMStation的2.9倍,而LLMStation仅为85%;与无额外硬件、保持100% SLO合规的vLLM+torchtune基线相比,DeltaServe的微调吞吐量也提升了39%。
英文摘要
LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a host-agnostic co-serving design that converts this idle inference capacity into LoRA fine-tuning throughput while preserving inference service-level objectives (SLOs). DeltaServe integrates with existing inference engines through a compact hook interface that requires only multi-LoRA batching support. It exploits the shared execution structure of inference prefill and LoRA fine-tuning forward passes, and uses an SLO-aware scheduler to admit and execute fine-tuning only when sufficient inference headroom is available. The scheduler is driven by a CUDA-graph-aware latency model calibrated offline and refined online. We integrate DeltaServe with vLLM, SGLang, and S-LoRA. On a production trace from Company X, DeltaServe on vLLM delivers 2.9x higher fine-tuning throughput than LLMStation at 100% inference SLO compliance, versus 85% for LLMStation. It also achieves 39% higher fine-tuning throughput than a baseline running vLLM+torchtune, using no additional hardware and maintaining full SLO compliance.