发表机构
University of California, Los Angeles; New York University; University of Utah(加州大学洛杉矶分校; 纽约大学; 犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对异构系统调度中故障影响,提出TFR-GNN图神经网络调度器,结合拓扑与故障感知,蒸馏容错预言机,显著降低预期完工时间并泛化至大规模工作流。
AI 中文摘要
在异构分布式系统上调度工作流有向无环图(DAG)是一个经典的NP难问题,列表调度启发式算法HEFT因其低复杂度和强完工时间而仍是事实上的标准。然而,在实际部署中,机器会发生故障:商用节点和可抢占节点远不如专用节点可靠,而一个完工时间最优但忽视可靠性的放置方案可能因节点故障而大幅减慢。我们在WfCommons/Pegasus语料库的真实工作流结构上实证表明,在故障强度和集群负载的联合空间中,没有单一固定启发式算法是最优的:无故障时HEFT最优,而在故障下,当存在备用容量时,可靠性感知放置可将预期完工时间最多降低52%。受此启发,我们提出TFR-GNN,一种图神经网络调度器,它结合了任务DAG上的双向依赖注意力、机器图(带宽加权)上的拓扑注意力,以及一个带有故障门控可靠性倾斜和可选复制门的交叉注意力放置头。我们通过将最佳组合容错预言机蒸馏为单一一次性策略来训练TFR-GNN。在真实工作流和双峰可靠性集群模型上,TFR-GNN在无故障时与HEFT完全匹配,在故障下将预期完工时间平均降低14.8%(最多47%),比固定可靠性感知基线(R-HEFT)提高11%,并作为单一策略匹配每场景事后预言机而无需任何部署时调优,且能泛化到未见过的应用和比训练时大一个数量级的工作流,同时为近5000个任务的图在远低于一秒内生成调度方案。所有结果均由经过验证的事件级模拟器在真实工作流数据上产生;没有实验数字是合成的。
英文摘要
Scheduling workflow directed acyclic graphs (DAGs) on heterogeneous distributed systems is a classical NP-hard problem, and the list-scheduling heuristic HEFT remains the defacto standard because of its low complexity and strong makespan. In real deployments, however, machines fail: commodity and pre-emptible nodes are far less reliable than dedicated ones, and a makespan-optimal but reliability-agnostic placement can be dramatically slowed by node failures. We show empirically, on real workflow structures from the WfCommons/Pegasus corpus, that no single fixed heuristic is best across the joint space of failure intensity and cluster load: with no failures HEFT is optimal, whereas under failures a reliability-aware placement can reduce the expected makespan by up to 52% when spare capacity exists. Motivated by this, we present TFR-GNN, a graph neural scheduler that combines bidirectional dependency attention over the task DAG, topology attention over the (bandwidth-weighted) machine graph, and a cross-attention placement head augmented with a failure-gated reliability tilt and an optional replication gate. We train TFR-GNN by distilling a best-of-portfolio fault-tolerant oracle into a single one-shot policy. On real workflows and a bimodal-reliability cluster model, TFR-GNN matches HEFT exactly when there are no failures, reduces the expected makespan under failures by 14.8% on average (up to 47%) over HEFT, beats a fixed reliability-aware baseline (R-HEFT) by 11% and matches a per-scenario hindsight oracle as a single policy without any deployment-time tuning, and generalises to unseen applications and to workflows an order of magnitude larger than those seen in training, while producing schedules in well under a second for graphs of nearly 5,000 tasks. All results are produced by a verified event-level simulator on real workflow data; no experimental numbers are synthetic.