LLMVisor:面向多租户大语言模型(LLM)服务的实时延迟归因模型
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
浏览论文内容
中文总结 AI 辅助
针对多租户LLM服务中联合批处理导致的延迟归因难题,提出基于屋顶线模型的LLMVisor实时延迟归因模型,在微秒级高效运行,相比基线方法大幅降低预填充和解码阶段的相对误差
中文摘要 AI 辅助
随着大语言模型(LLM)推理转向多租户GPU集群,联合批处理(co-batching)可提升吞吐量,但会模糊各租户的使用情况并限制管控。要实现推理引擎的分数级共享,需要一种准确且轻量、足以在调度循环内运行的实时逐请求归因原语。我们提出LLMVisor,一种基于屋顶线模型(roofline)的延迟归因模型,通过与浮点运算量(FLOPs)和内存I/O流量成比例的特征,以简洁的分段线性形式捕捉内存绑定和计算绑定阶段。LLMVisor将批处理延迟分解为可加的逐请求份额,且能在微秒级高效运行。我们在A100/H100 GPU上,针对Llama 3.1-8B、Qwen 2.5-14B/32B模型,在不同张量并行度和工作负载混合场景下对LLMVisor进行评估。与基于token计数的基线方法相比,尽管存在批处理变异性和序列差异,LLMVisor在预填充阶段的R平方值接近完美,且p90和p99分位的相对误差分别降低了最多2.5倍和3.3倍;在解码阶段则分别降低了最多3.5倍和4.4倍。
英文摘要
As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a real-time, per-request attribution primitive that is accurate and light enough to run inside the scheduling loop. We present LLMVisor, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic. LLMVisor decomposes batch latency into additive, per-request shares and runs efficiently at microsecond scale. We evaluate LLMVisor across Llama 3.1-8B and Qwen 2.5-14B/32B on A100/H100 GPUs under varying tensor parallelism and workload mixes. Compared to a token-count baseline, LLMVisor attains near-perfect R-squared and reduces relative error by up to 2.5x and 3.3x at p90 and p99, respectively, for prefill, and by up to 3.5x and 4.4x for decode, despite batching variability and sequence divergence.
发表机构
- University of Michigan(密歇根大学)
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。