arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22242cs.DC

用于异构人工智能工作负载的智能CPU-GPU调度

Agentic CPU-GPU Scheduling for Heterogeneous AI Workloads

Tianxi Lu, Sherief Reda

首次发表
浏览论文内容

中文总结 AI 辅助

研究异构人工智能工作负载的调度问题,提出将设备调度制定为三种选项,识别影响延迟的因素,构建智能调度器,该调度器在多种场景下达到最优映射,精度与经典基线匹配且无需离线训练,性能优于其他策略。

中文摘要 AI 辅助

智能人工智能系统在共享GPU/CPU基础设施上组合异构工具工作负载,但现有框架默认将所有具备GPU能力的工具分配到GPU上。我们对19种人工智能工具在GPU和CPU上进行分析,发现11种优先使用GPU,4种情况不明,1种因PCIe传输占优而优先使用CPU,3种对设备无偏好,这表明一概优先使用GPU的调度并非最优。我们将设备调度制定为在VRAM预算下为每个工具分配三种选项之一:立即在GPU上执行、排队在GPU上执行或卸载到CPU,并识别出导致端到端延迟与静态分析不同的两个运行时因素:GPU利用率竞争和VRAM容量竞争。我们提出了一种智能调度器,它将一个大语言模型智能体与一个算法运行时监视器配对,该监视器通过运行平均值、对称重新探测、交换重新探测和探索提示来扩展大语言模型能观察到的内容,而从不规定采用哪种映射。在涵盖串行执行、并行竞争和内存受限执行的13种场景中,该智能调度器在所有13种场景中都达到了暴力最优映射,在映射精度上与最佳经典基线匹配,同时避免了对完整映射进行类似强盗式的探索,并且在无需任何离线训练情况下优于HEFT、StarPU和全GPU策略。

英文摘要

Agentic AI systems compose heterogeneous tool workloads on shared GPU/CPU infrastructure, yet existing frameworks assign all GPU-capable tools to the GPU by default. We profile 19 AI tools across GPU and CPU and find that 11 are GPU-preferred, 4 are ambiguous, 1 is CPU-preferred due to PCIe transfer dominance, and 3 are device-neutral, establishing that blanket GPU-first scheduling is suboptimal. We formulate device scheduling as assigning each tool to one of three options: immediate GPU execution, queued GPU execution, or CPU offload, under a VRAM budget, and identify two runtime factors that cause end-to-end latency to diverge from static profiles: GPU utilization contention and VRAM capacity contention. We present an agentic scheduler that pairs an LLM agent with an algorithmic runtime monitor, where the monitor expands what the LLM can observe via running averages, symmetric reprobing, swap reprobing, and exploration hints, without ever prescribing which mapping to adopt. Across 13 scenarios spanning serial execution, parallel contention, and memory-constrained execution, the agentic scheduler reaches the brute-force optimal mapping in all 13 scenarios, matching the best classical baseline on mapping accuracy while avoiding bandit-style exploration over complete mappings, and outperforming HEFT, StarPU, and the all-GPU policy while requiring zero offline training.

补充信息

↑