ServerlessT2I:在无服务器平台上实现高效的文生图工作流服务
ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform
浏览论文内容
中文总结 AI 辅助
ServerlessT2I是将T2I工作流分解为独立模型函数的无服务器系统,通过优化调度与内存利用,在相同GPU预算下提升请求率2倍,固定请求率下节省GPU资源达3倍。
中文摘要 AI 辅助
文生图(T2I)工作流越来越多地被部署在无服务器平台上,因为用户经常会组合自定义工作流并间歇性地调用它们。现有平台通常将每个工作流部署为一个不透明的GPU函数,将工作流中的所有组成模型一起置备、放置和扩展。这种单体设计会模糊工作流结构、增加扩展开销、迫使用户管理底层GPU协调,并限制多租户集群中的细粒度公平性。在本文中,我们提出了ServerlessT2I,这是一个原生无服务器系统,它将T2I工作流分解为松耦合的模型函数,这些函数可以被独立管理和调度。通过显式管理单个模型的执行,ServerlessT2I实现了按模型扩展、声明式工作流组合、透明的GPU驻留通信以及感知公平性的调度。为了使这种分解高效,ServerlessT2I利用了计算密集型T2I推理所闲置的GPU内存,以构建一个数据平面,从而减少模型加载和数据通信开销。该系统还为多租户服务引入了一个公平调度器。使用生产追踪数据,ServerlessT2I在相同的GPU预算下,能维持比现有T2I工作流服务系统高2倍的请求率;在固定请求率下,它能节省多达3倍的GPU资源,同时满足服务水平目标(SLO)。
英文摘要
Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together. This monolithic design obscures workflow structure, inflates scaling overhead, forces users to manage low-level GPU coordination, and limits fine-grained fairness in multi-tenant clusters. In this paper, we present ServerlessT2I, a serverless-native system that decomposes a T2I workflow into loosely coupled model functions that can be independently managed and scheduled. By explicitly managing individual model execution, ServerlessT2I enables per-model scaling, declarative workflow composition, transparent GPU-resident communication, and fairness-aware scheduling. To make this decomposition efficient, ServerlessT2I harvests slack GPU memory left idle by compute-bound T2I inference to build a data plane that reduces model loading and data communication overheads. \sys{} further introduces a fair scheduler for multi-tenant serving. Using production traces, ServerlessT2I sustains up to 2$\times$ higher request rates than existing T2I workflow serving systems with the same GPU budget; for a fixed request rate, it saves up to 3$\times$ GPU resources while satisfying service level objectives (SLOs).
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。