arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mosaic:通过内核级干扰预测实现具有延迟保证的GPU共享

Mosaic: GPU Sharing with Latency Guarantees through Kernel-Level Interference Prediction

Foteini Strati, Ethan Graham, Leo Stephan, Paul Elvinger, Ana Klimovic

arXiv 2610.07504首次发表:更新:

发表机构

ETH Zurich; University of Oxford; Wayve(苏黎世联邦理工学院; 牛津大学; Wayve)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Mosaic通过内核级干扰预测模型显式建模GPU干扰机制,结合分析与轻量学习,降低预测误差达一个数量级,并集成调度器实现延迟保证与吞吐量最大化。

AI 中文摘要

GPU在AI工作负载中的需求日益增长,但往往仍处于大幅未充分利用的状态,这促使了工作负载共置。然而,共置会引入干扰,可能降低对延迟关键型工作负载的性能。现有方法使用基于启发式的调度或干扰预测器来缓解干扰。启发式方法依赖于粗粒度指标,忽视了重要的干扰来源,而预测器通常依赖于模拟器特定或同样粗粒度的指标。然而,GPU干扰是复杂的,源于多种机制,包括线程块放置、内存层次结构争用和SM内资源争用。我们提出了Mosaic,一种内核级干扰预测器,通过结合分析和轻量级学习模型显式建模这些机制。在四种GPU架构上,与之前的预测器相比,Mosaic将预测误差降低了最多一个数量级。我们将Mosaic集成到一个调度器MosaicSched中,该调度器执行在线内核准入控制,并在全GPU共置和SM分区之间进行选择,以在满足延迟SLO的同时最大化尽力而为的吞吐量。在所有工作负载中,MosaicSched将p99延迟保持在目标SLO以下或非常接近目标SLO。

英文摘要

GPUs are increasingly in demand for AI workloads, yet often remain substantially underutilized, motivating workload colocation. However, colocation introduces interference that can degrade latency-critical workloads. Existing approaches mitigate interference using either heuristic-based scheduling or interference predictors. Heuristics rely on coarse-grained metrics that overlook important interference sources, while predictors often depend on simulator-specific or similarly coarse metrics. However, GPU interference is complex and arises from multiple mechanisms, including thread-block placement, memory hierarchy contention, and intra-SM resource contention. We present Mosaic, a kernel-level interference predictor that explicitly models these mechanisms using a combination of analytical and lightweight learned models. Across four GPU architectures, Mosaic reduces prediction error by up to an order of magnitude compared to prior predictors. We integrate Mosaic into a scheduler, MosaicSched, that performs online kernel admission control and selects between full-GPU colocation and SM partitioning to maximize best-effort throughput while satisfying latency SLOs. Across all workloads, MosaicSched keeps the p99 latency below or very close to the target SLO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑