arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图注意力驱动的分层强化学习的云工作流调度

Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement Learning

Zongjin Li, Shaohan Feng, Chunxi Yang, Wenbo Wang

arXiv 2609.14952首次发表:更新:

发表机构

Kunming University of Science and Technology; Zhejiang Gongshang University(昆明理工大学; 浙江工商大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出图注意力驱动的分层强化学习框架,通过图注意力网络提取DAG依赖信息,结合事件驱动SMDP和双PPO代理,在阿里巴巴集群轨迹上实现高成功率、高容器利用率和低能耗的云工作流调度。

AI 中文摘要

动态云工作流调度必须在满足截止期限、容器利用率和能耗之间取得平衡,同时处理随机任务执行速度、依赖放置的通信以及任务与容器决策的耦合。工作流自然建模为有向无环图(DAG),但传统的基于向量或矩阵的状态不能完全捕获其依赖拓扑。为了更好地表示任务紧迫性和结构关系,我们为任务分配预测的子截止期限,并使用多头图注意力网络(GAT)从演化的DAG中提取依赖信息。基于这些表示,我们开发了图注意力驱动的分层强化学习(GA-HRL)框架,并将调度过程建模为事件驱动的分层半马尔可夫决策过程(SMDP)。工作流到达和任务完成触发调度事件。在每个调度事件中,任务调度(TS)代理首先处理当前就绪的任务,将它们分配给可用的现有容器或请求新容器。请求的容器随后由容器调度(CS)代理在环境推进之前进行主机放置。两个代理使用独立的近端策略优化(PPO)交替训练。在2018年阿里巴巴集群轨迹上的实验表明,GA-HRL保持了有竞争力的工作流成功率,并且在成功率相当的情况下,通常实现更高的容器利用率和更低的能耗。在最大的速度变化下,它以较小的成功率差距换取显著更低的能耗。仿真代码可在以下网址获取:此https URL。

英文摘要

Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-execution speeds, placement-dependent communication, and coupled task and container decisions. Workflows are naturally modeled as directed acyclic graphs (DAGs), but conventional vector- or matrix-based states do not fully capture their dependency topology. To better represent task urgency and structural relationships, we assign predicted sub-deadlines to tasks and use a multi-head graph attention network (GAT) to extract dependency information from the evolving DAGs. Based on these representations, we develop a Graph Attention-Driven Hierarchical Reinforcement Learning (GA-HRL) framework and model the scheduling process as an event-driven hierarchical semi-Markov decision process (SMDP). Workflow arrivals and task completions trigger scheduling events. At each scheduling event, the Task Scheduling (TS) agent first processes the currently ready tasks by assigning them to admissible existing containers or requesting new ones. The requested containers are then processed by the Container Scheduling (CS) agent for host placement before the environment advances. The two agents are trained alternately using separate Proximal Policy Optimization (PPO). Experiments on the 2018 Alibaba cluster trace show that GA-HRL maintains competitive workflow success rate and, in settings where success is comparable, generally achieves higher container utilization and lower energy consumption. Under the largest speed variation, it trades a small success-rate margin for substantially lower energy. Simulation code is available at: https://github.com/zongjin130/GA-HRL.

CommentsPaper submitted to IEEE Internet of Things Journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑