基于资源利用率的Spark应用能耗估算精度如何?
How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation?
- University of Glasgow(格拉斯哥大学)
- Humboldt-Universität zu Berlin(柏林洪堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究以Kubernetes上的Apache Spark为对象,对比模型能耗估算与Intel RAPL实测能耗,发现外部监测可降低Spark应用的封装级能耗低估误差,为云环境下Spark应用能耗估算提供了优化方向。
AI中文摘要:
分布式批量数据处理应用广泛运行于云资源上,而用户对节点级硬件能耗计数器的访问受限,阻碍了透明的可持续性核算。因此能耗与碳归因方法依赖于功率模型和可用资源利用率轨迹,但在有直接计数器可用时,需验证这些估算的准确性。本研究以在Kubernetes上运行的Apache Spark作为案例研究的数据流运行时和集群资源管理器,在AWS裸金属云和本地集群上,将基于模型的能耗估算与Intel RAPL封装级及DRAM能耗进行比较,对比不同CPU使用信号和内存系数。结果显示,与Spark任务轨迹相比,外部监测可降低有符号封装级能耗误差,AWS上的低估从-29.58%降至-24.41%,本地集群上的低估从-24.00%降至-16.22%。
英文摘要:
Distributed batch data processing applications are widely executed on cloud-based resources where restricted user access to node-level hardware energy counters hinders transparent sustainability accounting. Energy and carbon attribution methodologies therefore depend on power models and available resource utilisation traces, yet the accuracy of these estimates has to be validated while direct counters are available. In this work, we use Apache Spark running on Kubernetes as a case-study dataflow runtime and cluster resource manager to compare model-based energy estimates to Intel RAPL package and DRAM energy on an AWS bare-metal cloud and an on-premises cluster, comparing different CPU usage signals and memory coefficients. We show that external monitoring improves signed package-energy error relative to Spark task traces, reducing underestimation from -29.58% to -24.41% on AWS and from -24.00% to -16.22% on-premises.