arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10956cs.DC

ClusterBench:面向集群范围的持续基准测试与回归测试框架

ClusterBench: A Framework for Cluster-Wide Continuous Benchmarking and Regression Testing

Aditya Ujeniya, Jan Eitzinger, Thomas Gruber, Georg Hager, Gerhard Wellein

首次发表
浏览论文内容

中文总结 AI 辅助

ClusterBench是面向集群范围的持续基准测试框架,含多组件基准测试集合,可检测软件变更引发的性能回归,其在NHR@FAU集群的测试揭示了硬件变异性及冷却方式对性能的影响。

中文摘要 AI 辅助

数据中心需要能验证整个部署而非单个节点的工具,用于验收阶段及之后的定期检查。这要求在单次提交中将相同基准测试分发到集群的每个节点,因此需要具备集群感知的调度能力。本文提出ClusterBench,一个面向集群范围的持续基准测试框架,它附带针对CPU、GPU、内存、互连和I/O各组件的基准测试集合。由于测量会在集群的整个生命周期内重复进行,ClusterBench会收集跨空间和时间的数据。将测量结果与早期运行对比,可检测内核更新或新库版本等软件变更引入的性能回归。这些测量还构成了用于硬件变异性研究的数据集。在NHR@FAU的Helma、Alex和Fritz集群上,单个组件内的变异性保持在1%以内;尽管节点规格相同,不同样本间的变异性达到5%。将性能与功耗、频率和温度关联分析显示,风冷节点与液冷节点的该关系存在差异。

英文摘要

Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for cluster-wide continuous benchmarking. It ships with a benchmark collection targeting each component: CPU, GPU, memory, interconnect, and I/O. Because measurements are repeated throughout the cluster's lifetime, ClusterBench collects data across space and time. Comparison against earlier runs detects performance regressions introduced by software changes, such as kernel updates or new library versions. The measurements also form a dataset for research on hardware variability. On the NHR@FAU clusters Helma, Alex, and Fritz, variation within a single component stays within 1%. Variation across specimens reaches 5%, despite nodes identical by specification. Correlating performance with power draw, frequency, and temperature shows that this relationship differs between air- and liquid-cooled nodes.

↑