分解预测性Kubernetes自动扩缩容以应对长启动延迟下的大语言模型服务
Decomposing Predictive Kubernetes Autoscaling for Large Language Model Serving Under Long Startup Delays
- University of California, San Diego(加州大学圣地亚哥分校)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究分解预测性Kubernetes自动扩缩容组件,发现EWMA预测器结合延迟前瞻和UCB边际可大幅降低LLM服务的TTFT SLO违规率,前瞻是最大贡献因素。
AI中文摘要:
部署在Kubernetes上的大语言模型(LLM)推理面临传统Web服务所不具备的自动扩缩容挑战:由于必须加载数GB的模型权重,新的服务副本需要两到十分钟才能启动,这使得纯反应式扩缩容在结构上滞后。我们提出一个尖锐的问题:在预测性自动扩缩容器的各个组件中,在如此长的执行延迟下,哪些组件真正重要?我们通过将预测性自动扩缩容分解为四个因素——令牌感知的需求跟踪、启动延迟前瞻、有界不确定性边际和对象状态观测——并在由ServeGen生成的生产衍生重尾工作负载上分别测量每个因素的贡献来回答这个问题。我们的主要发现是,一个简单的指数加权移动平均(EWMA)预测器,结合延迟感知的前瞻和上置信界(UCB)边际,捕获了大部分收益,在五个随机种子中将首令牌时间(TTFT)服务级别目标(SLO)违规率从53%(反应式,基于每秒查询数)降至0.5%;仅前瞻一项就是最大的单一因素,减少了14倍。卡尔曼滤波变体并未一致地改善成本-SLO权衡。受控实验隔离了令牌粒度必要性的原因:上下文长度通过键值缓存压力,在匹配吞吐量下对TTFT的损害远大于请求速率。最后,在真实Kubernetes集群(Qwen2.5-7B,A100,vLLM)上的验证确认了核心机制:延迟感知的前瞻控制器将TTFT违规率相对于反应式KEDA扩缩容从63.5%降至3.7%。我们始终区分哪些发现特定于Kubernetes执行,哪些普遍适用于LLM服务。
英文摘要:
Large language model (LLM) inference deployed on Kubernetes faces an autoscaling challenge that conventional web services do not: new serving replicas take two to ten minutes to start because multi-gigabyte model weights must be loaded, which makes purely reactive scaling structurally late. We ask a sharp question: among the components of a predictive autoscaler, which ones actually matter under such long actuation delays? We answer it by decomposing predictive autoscaling into four factors---token-aware demand tracking, startup-delay lookahead, a bounded uncertainty margin, and plant-state observation---and measuring each factor's contribution in isolation on production-derived heavy-tailed workloads generated by ServeGen. Our main finding is that a simple exponentially weighted moving average (EWMA) predictor with delay-aware lookahead and an upper confidence bound (UCB) margin captures most of the benefit, reducing time-to-first-token (TTFT) service-level-objective (SLO) violations from 53% (reactive, queries-per-second based) to 0.5% across five random seeds; lookahead alone is the single largest factor, a 14$\times$ reduction. Kalman filter variants do not consistently improve the cost--SLO tradeoff. Controlled experiments isolate the reason token granularity is necessary: context length, through key--value cache pressure, degrades TTFT far more than request rate at matched throughput. Finally, a validation on a real Kubernetes cluster (Qwen2.5-7B, A100, vLLM) confirms the central mechanism: a delay-aware lookahead controller cuts TTFT violations from 63.5\% to 3.7\% relative to reactive KEDA scaling. We distinguish throughout which findings are specific to Kubernetes actuation and which are general to LLM serving.