链路中断下DAG工作流的调度修复
Schedule Repair for DAG Workflows under Link Disruptions
浏览论文内容
中文总结 AI 辅助
本研究针对DAG工作流在链路中断下的调度修复问题,提出多种修复策略并评估其性能,发现修复范围应适应中断过程与成本,而非固定不变。
中文摘要 AI 辅助
在联网物联网系统中,有向无环图(DAG)工作流的调度通常是在假设网络静态或一般稳定的情况下计算的。在对抗性和敌意环境中,这一假设不成立。由于移动性、干扰和干扰,链路会降级和失效。我们研究调度修复:当链路中断使部分调度失效时,应重新调度多少部分?我们引入了一系列修复策略,这些策略在修复范围上有所不同,即每个策略可能移动多少待处理调度:等待中断结束、将数据绕道传输、仅局部重新调度受影响的任务,或全局重新调度所有待处理任务。我们针对每个策略与一个神谕进行评估,并为每次修复收取与移动程度成比例的决策延迟。在跨越合成任务图、RIoTBench管道和WfCommons科学工作流的100个工作负载实例中,每个实例在五个通信与计算比率(CCRs)下运行,并由具有故意不同相关结构的过程中断,我们发现没有单一范围获胜:绕道几乎消除了等待30%代价的孤立故障;在干扰停电下,全局修复在神谕的4%以内;在无记忆链路抖动下,对于较大和通信密集型工作负载,等待受到青睐(这是路由抖动阻尼的调度类比);自愈移动性中断奖励耐心而非反应;考虑修复延迟首先侵蚀大范围。我们得出结论,修复的范围应适应中断过程和修复成本,而不是由调度器固定。
英文摘要
Schedules for directed acyclic graph (DAG) workflows in networked IoT systems are typically computed assuming a static or generally stable network. In contested and adversarial environments, this assumption is not valid. Links degrade and fail due to mobility, interference, and jamming. We study schedule repair: when a link disruption invalidates part of a schedule, how much of it should be rescheduled? We introduce a spectrum of repair policies that vary in repair scope, how much of the pending schedule each may move: wait out the disruption, reroute data around it, reschedule only the affected tasks locally, or reschedule all pending tasks globally. We evaluate each against an oracle and charge every repair a decision latency proportional to the extent to which it moves. Across 100 workload instances spanning synthetic task graphs, RIoTBench pipelines, and WfCommons scientific workflows, each run at five communication-to-computation ratios (CCRs) and disrupted by processes with deliberately different correlation structure, we find that no single scope wins: rerouting nearly erases isolated failures that cost waiting 30%, global repair comes within 4% of the oracle under jamming blackouts, waiting is favored under memoryless link flapping for larger and communication-heavy workloads (the scheduling analog of route-flap damping), self-healing mobility outages reward patience over reaction, and accounting for repair latency erodes large scopes first. We conclude that the scope of the repair should be adapted to the disruption process and the repair cost, rather than fixed by the scheduler.
发表机构
- University of Southern California(南加州大学)
- Loyola Marymount University(洛约拉马利蒙特大学)
- U.S. DEVCOM Army Research Laboratory(美国DEVCOM陆军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。