用于共享GPU上并发异构AI推理的可扩展运行时调度的MeanField代理模型
MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
浏览论文内容
中文总结 AI 辅助
针对共享GPU上并发异构AI推理的调度难题,提出MeanField代理模型,将其集成到遗传算法调度器中,在多类工作负载场景下实现高精度、低耗时的可扩展调度,性能接近穷举搜索且效率大幅提升。
中文摘要 AI 辅助
将异构AI模型并发部署在共享GPU上会引发资源竞争,使运行时调度变得复杂。代理模型可避免高昂的在线基准测试成本,但其分析需求通常随并发运行模型数量呈组合式增长,限制了可扩展性。本文提出一种MeanField代理模型,该模型基于局部配置和GPU聚合状态预测各模型性能,而非显式建模所有联合交互。针对N∈{2,3,4,5,6}的并发大语言模型(LLM)与视觉工作负载的实验显示,其预测精度较高(R²≈0.96),经验样本预算随N呈近似线性增长,与完全联合分析的组合式成本形成对比。将该代理模型集成到遗传算法调度器中,可扩展至含78732个可行联合配置的N=5问题,在8种动态工作负载场景下,其结果与穷举搜索的偏差保持在0.10%以内,且无服务水平协议(SLA)弃权(不执行)情况;完整在线遗传算法决策的中位数耗时为26毫秒,比穷举代理搜索快约5倍。
英文摘要
Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scalability. We propose a MeanField surrogate that predicts per-model performance from local configuration and aggregate GPU state rather than explicitly modeling all joint interactions. Experiments on concurrent LLM and vision workloads across $N \in \{2,3,4,5,6\}$ show high predictive accuracy ($R^2 \approx 0.96$) with an empirical sample budget that grows approximately linearly in $N$, in contrast to the combinatorial cost of fully joint profiling. Integrated into a genetic algorithm scheduler, the surrogate scales to an $N=5$ problem with 78,732 feasible joint configurations, remaining within 0.10% of the exhaustive search with zero SLA violations across eight dynamic workload scenarios, while complete online GA decisions take 26 ms median, about $5\times$ faster than exhaustive surrogate search.
发表机构
- Seoul National University(首尔大学)
- Institute of Computer Technology (ICT), Seoul National University(首尔大学计算机技术研究所)
机构由 AI 辅助整理,请以论文原文为准。