发表机构
Beihang University; University of Leeds(北京航空航天大学; 利兹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对异构GPU上的多智能体大语言模型工作流,提出预测引导的运行时系统,通过优化编排降低了端到端延迟与GPU资源消耗。
AI 中文摘要
并发多智能体工作流在异构GPU池上运行时,会暴露未来依赖关系和服务状态需求,而GPU池存在时变负载、模型驻留情况和资源可用性。逻辑工作流定义了所需计算,但其物理调度单元、模型生命周期操作、资源排序和放置必须根据观测到的池状态选择。我们提出一种预测引导的运行时,利用工作流预测构建并优化物理执行图。预测器估计设备特定的激活延迟、峰值内存和模型加载成本,再通过工作流依赖关系传播这些预测,以预测激活就绪情况和未来模型需求。构造器构建语义保留的融合与模型生命周期备选方案,调度器则根据实时池状态联合优化它们的选择、放置和执行顺序。在包含异构GPU池上三个工作流场景的工作负载中,与最先进的工作流调度器相比,我们的系统在突发到达下将端到端总时间和整体p95完成延迟分别降低了高达36.8%和25.9%,还为每个完成的会话节省了高达24.63 GPU-秒。
英文摘要
Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, resource ordering, and placement must be selected according to the observed pool state. We present a prediction-guided runtime that uses workflow forecasts to construct and optimize a physical execution graph. Predictor estimates device-specific activation latency, peak memory, and model-loading cost, then propagates these predictions through workflow dependencies to forecast activation readiness and future model demand. Constructor builds semantics-preserving fusion and model-lifecycle alternatives, while Scheduler jointly optimizes their selection, placement, and execution order based on the live pool state. Across a workload spanning three workflow scenarios on a heterogeneous GPU pool, our system reduces end-to-end makespan and overall p95 completion latency under burst arrivals by up to 36.8% and 25.9%, respectively, over state-of-the-art workflow schedulers. It also saves up to 24.63 GPU-s per completed session.
Comments13 pages, 7 figures