AI 中文总结
HorizonServe是一款单GPU全模态模型服务系统,通过协调请求准入与GPU分配,提升了SLO达成率并降低了首响应延迟,解决了全模态模型统一部署的调度问题。
AI 中文摘要
全模态模型在单一服务后端中统一了文本、语音、图像及多模态推理能力,但这种统一部署带来了新的调度问题。不同输出模态的请求可共享初始多模态骨干网络,随后分化至下游生成阶段,导致同一GPU上出现异构的首响应指标与服务水平目标(SLO)。现有大语言模型(LLM)及多模态服务系统主要优化令牌进度或输入侧处理,未联合控制共享阶段的时间共享与并发运行阶段的空间共享。本文提出HorizonServe,一款单GPU全模态模型服务系统,用于在异构SLO下协调请求准入与GPU分配。HorizonServe会分析各类别首响应延迟、保护松弛度有限的请求、在各执行路径间轮换共享阶段的执行机会,并在下游阶段活跃时限制共享阶段的流式多处理器(SM)分配。在三种全模态模型工作负载与两种GPU平台上,HorizonServe在到达率扫描中将SLO达成率提升了最高4.9倍,在下游密集流量下提升了7.0倍,并将各类别首响应延迟降低了38.4%至63.7%。
英文摘要
Omni models unify text, speech, image, and multimodal reasoning in a single serving backend, but this unified deployment exposes a new scheduling problem. Requests with different output modalities may share an initial multimodal backbone and then diverge into downstream generation stages, creating heterogeneous first-response metrics and service-level objective (SLO) targets on the same GPU. Existing large language model (LLM) and multimodal serving systems mainly optimize token progress or input-side processing, and they do not jointly control temporal sharing in the shared stage and spatial sharing among co-running stages. This paper presents HorizonServe, a single-GPU omni-model serving system that coordinates request admission and GPU allocation under heterogeneous SLOs. HorizonServe profiles per-class first-response latency, protects requests with limited slack, rotates shared-stage opportunities across execution paths, and throttles the shared-stage streaming multiprocessor (SM) allocation when downstream stages are active. Across three omni-model workloads and two GPU platforms, HorizonServe improves SLO attainment by up to 4.9$\times$ in arrival-rate sweeps and 7.0$\times$ under downstream-heavy traffic, and reduces per-class first-response latency by 38.4--63.7\%.