AI 中文总结
该研究针对云集群调度的SLO过度严格执行与共享资源拥塞感知缺失问题,提出软SLO限制与感知资源的调度策略,可降低重调度动作49%、节点拥塞8%,提升调度效率与性能保障。
AI 中文摘要
云环境中的工作负载调度通常依赖于对应用程序资源需求和硬件利用率的简单假设,忽视了应用层面的性能目标和硬件资源争用,这会导致资源使用效率低下和性能下降。本文针对当前方法的两个关键局限展开研究:其一,对服务水平目标(SLO)的过度严格执行往往会导致资源利用率低下和能效不佳;其二,末级缓存(LLC)、内存带宽等共享资源缺乏拥塞感知。为此,本文提出两种互补策略以解决上述局限:(i)整合软SLO限制,允许受控的资源过度分配并容忍轻微、瞬时的违规,以提升集群效率;(ii)基于末级缓存(LLC)、内存带宽等共享资源的实时拥塞洞察,引入感知资源的调度与重调度。实验结果显示,与硬限制方法相比,软SLO限制可将 corrective rescheduling actions 减少49%,同时维持可接受的性能保障;此外,感知资源的调度可将节点级拥塞降低8%,并进一步缓解SLO违规,证明了在调度与重调度决策中纳入应用层面灵活性和硬件层面洞察的有效性。
英文摘要
Workload scheduling in cloud environments often relies on simplistic assumptions about application resource needs and hardware utilization. Overlooking application-level performance objectives and hardware resource contention that leads to inefficient resource usage and degraded performance. This paper addresses two key limitations of current approaches. First, unnecessarily strict enforcement of service level objectives (SLOs) often leads to resource underutilization and poor energy efficiency. Second, lack of congestion awareness in shared resources such as last-level cache (LLC) and memory bandwidth. In this paper, we propose two complementary strategies to address these limitations: (i) integrating soft SLO limits that allow controlled overcommitment and tolerate minor, transient violations to improve cluster efficiency, and (ii) introducing resource-aware scheduling and rescheduling based on real-time congestion insights for shared resources such as last-level cache (LLC) and memory bandwidth. Our results show that soft SLO limits reduce corrective rescheduling actions by 49% compared to hard-limit approaches while maintaining acceptable performance guarantees. Additionally, resource-aware scheduling decreases node-level congestion by 8% and further mitigates SLO violations, demonstrating the effectiveness of incorporating application-level flexibility and hardware-level insights into scheduling and rescheduling decisions.
Comments10 pages, 9 figures, 3 tables