arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HPC-MQBench:基于单代理Kafka评估的Slurm优先资格基准测试

HPC-MQBench: Qualification-First Benchmarking on Slurm with a Single-Broker Kafka Evaluation

Sepehr Mahmoodian, Julian Kunkel

arXiv 2610.09786首次发表:更新:

发表机构

University of Göttingen; Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG)(哥廷根大学; 哥廷根科学数据处理有限公司(GWDG))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HPC-MQBench,一个基于Slurm和单代理Kafka的基准测试框架,通过资格优先评估和可恢复控制,实现分布式实验控制与可审计配置选择,实验验证了其有效性。

AI 中文摘要

在高性能计算集群上进行消息传递实验,必须在调度器分配内协调服务、客户端和测量。我们提出了HPC-MQBench,一个由Slurm编排的基准测试,它将可恢复的实验控制与记录检查、交付核算和速率排名前的资格认证联系起来。其Kafka评估将生产者/控制器、一个代理、消费者和监控分布在四个节点上。使用内存支持的日志,120个工作负载配置产生了99个合格观测,14个超出生产者交付策略,7个证据无效。两个验证阶段各重复了五次块中的十个工作负载。在最后阶段,选定的工作负载在所有五次观测中均合格,平衡端点速率的中位数为每秒2927兆字节,待交付百分比为1.24%,第99百分位延迟为2.56秒。在早期阶段,它仅在两次观测中合格,在分配边界后资格有所改善,但未建立因果分配效应。选择取决于阶段和交付策略。资源测量未隔离出唯一瓶颈。贡献在于为分布式实验控制和在固定代理数量下的可审计配置选择提供了一个框架;多代理支持和扩展实验仍是未来工作。

英文摘要

Messaging experiments on high-performance computing clusters must coordinate services, clients, and measurement within scheduler allocations. We present HPC-MQBench, a Slurm-orchestrated benchmark that links resumable experiment control to record checks, delivery accounting, and qualification before rate ranking. Its Kafka evaluation placed producers/controller, one broker, consumers, and monitoring on four nodes. With memory-backed logs, 120 workload configurations yielded 99 qualified observations, 14 outside the producer-delivery policy, and seven with invalid evidence. Two validation stages each repeated ten workloads in five blocks. In the final stage, the selected workload qualified in all five observations, with medians of 2,927 mebibytes per second balanced endpoint rate, 1.24 percent pending deliveries, and 2.56 seconds for the 99th-percentile latency. It qualified in only two observations in the earlier stage, where qualification improved after an allocation boundary without establishing a causal allocation effect. Selection is conditional on the stage and delivery policy. Resource measurements did not isolate a unique bottleneck. The contribution is a framework for distributed experiment control and auditable configuration selection at a fixed broker count; multi-broker support and scaling experiments remain future work.

Comments8 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑