发表机构
Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HiGFRL提出层级图融合强化学习框架,通过三级状态表示与双网络架构,在异构云中实现依赖感知任务调度,显著降低Makespan并提升QoS。
AI 中文摘要
在异构云集群中,对具有依赖关系的任务进行在线调度是一个基础且具有挑战性的问题,这源于DAG拓扑结构与多维资源约束之间的复杂交互。尽管深度强化学习(DRL)已展现出潜力,但现有的基于图神经网络(GNN)的方法往往难以高效建模高阶拓扑依赖,并且任务状态与资源状态之间的耦合松散,导致调度决策短视。为解决这些局限,我们提出了HiGFRL,一种层级图融合驱动的强化学习框架。HiGFRL构建了一种新颖的三级状态表示,包括静态超图、动态全局图和局部二分图,以显式建模任务依赖与实时集群动态之间的交互。具体而言,我们设计了一种融合驱动的双网络架构来优化强化学习决策,其中上下文融合分配器将局部二分匹配特征与融合的全局上下文相结合,以执行精确的任务到节点分配;全局状态评估器利用全局动态图表示来准确估计预期的长期累积奖励。此外,我们引入了一种拓扑先验引导的混合奖励机制,将静态拓扑先验提炼到学习过程中以加速收敛。使用真实阿里巴巴集群轨迹进行的大量实验表明,HiGFRL显著优于启发式方法和DRL基线。具体而言,在具有挑战性的大规模高负载场景中,HiGFRL将Makespan降低了高达32.55%,并将平均任务流时间和平均任务等待时间分别优化了13.58%和13.79%。实验结果证实,HiGFRL不仅显著提高了集群吞吐量,还通过大幅减少排队延迟确保了卓越的服务质量(QoS)。代码发布:此https URL。
英文摘要
Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay between DAG topologies and multi-dimensional resource constraints. While DRL has shown promise, existing GNN-based approaches often struggle to efficiently model high-order topological dependencies and suffer from loose coupling between task and resource states, leading to myopic scheduling decisions. To address these limitations, we propose HiGFRL, a Hierarchical Graph Fusion-Driven Reinforcement Learning framework. HiGFRL constructs a novel three-level state representation comprising a Static Hypergraph, a Dynamic Global Graph, and a Local Bipartite Graph to explicitly model the interplay between task dependencies and real-time cluster dynamics. Specifically, we design a fusion-driven dual-network architecture to optimize RL decision-making, where a Context Fusion Allocator integrates local bipartite matching features with fused global context to execute precise task-to-node allocation, and a Global State Evaluator leverages the global dynamic graph representation to accurately estimate expected long-term cumulative reward. Furthermore, we incorporate a topology-prior-guided hybrid reward mechanism that distills static topological priors into the learning process to accelerate convergence. Extensive experiments using real-world Alibaba cluster traces demonstrate that HiGFRL significantly outperforms heuristics and DRL baselines. Specifically, in challenging large-scale high-load scenarios, HiGFRL reduces the Makespan by up to 32.55%, and optimizes the average task flow time and average task wait time by 13.58% and 13.79%, respectively. Experimental results confirm that HiGFRL not only significantly improves cluster throughput but also ensures superior QoS by substantially reducing queuing delays. Code Release:https://github.com/igeng/HiGFRL.
Comments16 pages, 13 figures