arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Roomie:用于高效模型服务的干扰感知共置

Roomie: Interference-Aware Colocation for Efficient Model Serving

Youssouph Faye, Francescomaria Faticanti, Shubham Jain, Francesco Bronzino

arXiv 2607.16784首次发表:更新:

AI 中文总结

针对DNN推理中GPU容量供不应求需共置模型的情况,Roomie架构通过解耦离线内核分析与在线干扰预测,用基于占用的模型预测干扰,经启发式算法和在线放置算法,有效降低推理延迟,提升吞吐量。

AI 中文摘要

随着对DNN推理需求的增长,GPU容量越来越供不应求,促使运营商在云及边缘部署中于同一设备上共置多个模型。共置是否成功或违反服务水平目标(SLO)取决于并发执行模型内核的时间重叠,现有服务系统要么忽略此效应,要么用无法捕捉时间动态的聚合资源配置文件进行近似处理。本文提出Roomie,一种模型服务编排架构,可预测并避免共置DNN之间的内核级干扰。Roomie将离线内核分析与在线干扰预测解耦,仅用分析提取每个内核的资源配置,用基于占用的分析模型预测干扰,该模型不受分析器引起的定时失真影响。然后,成对贪婪启发式算法以多项式而非指数时间近似多模型干扰,在线放置算法利用这些估计将每个传入模型分配到使预测减速最小的GPU上。我们的实验评估将Roomie与云级服务器集群和嵌入式边缘设备上的最先进解决方案进行比较,表明Roomie将SLO违规(即推理延迟)降低多达3倍,同时相对于现有方法保持相当且在许多情况下更优的吞吐量。

英文摘要

As demand for DNN inference grows, GPU capacity is increasingly oversubscribed, forcing operators to colocate multiple models on the same device in both cloud and edge deployments. Whether colocation succeeds or violates SLOs depends on the temporal overlap of kernels from concurrently executing models -- an effect that existing serving systems either ignore or approximate using aggregate resource profiles that fail to capture temporal dynamics. This paper presents Roomie, a model serving orchestration architecture that predicts and avoids kernel-level interference between colocated DNNs. Roomie decouples offline kernel profiling from online interference prediction. It uses profiling only to extract per-kernel resource configurations, and predicts interference with an occupancy-based analytical model immune to profiler-induced timing distortion. A pairwise greedy heuristic then approximates multi-model interference in polynomial rather than exponential time, and an online placement algorithm then uses these estimates to assign each incoming model to the GPU that minimizes predicted slowdown. Our experimental evaluation compares Roomie against state-of-the-art solutions across both cloud-grade server clusters and embedded edge devices, demonstrating that Roomie reduces SLO violations (i.e., inference latency) by up to 3x, while maintaining comparable, and in many cases superior, goodput relative to existing approaches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑