平衡工作负载性能与Slurm压力:四种Nextflow部署策略
Balancing Workload Performance and Slurm Stress: Four Nextflow Deployment Strategies
浏览论文内容
中文总结 AI 辅助
本文提出可复现测量协议,在ASU Phoenix和Dev集群评估四种Nextflow部署策略,发现存在时间与RPC的权衡,助力HPC站点选择兼顾性能与调度器压力的部署方案。
中文摘要 AI 辅助
在共享Slurm集群上,Wide Nextflow的扇出可提交数万个短任务。部署选择(单个作业、作业数组或分配内的嵌套调度器)会同时影响工作流周转时间和RPC(远程过程调用)量,RPC作为共享成本可能降低调度器响应速度。现有研究对比的是整个工作流系统,而按任务划分的排队指标未涵盖在现有分配内进行调度的架构。本文提出一种可复现的测量协议和基准测试工具,包含两个关键要素:一是在任何后端服务或分配请求前启动干净时钟,将不同架构置于同一时间轴上;二是通过每用户Slurm sdiag计数器测量可归因的RPC需求,将控制器处理时间作为独立于集群全局状态的敏感性上下文进行报告。我们在共享ASU Phoenix生产集群上评估四种多节点策略:Slurm原生、Slurm作业数组、HyperQueue和Flux;在单用户Dev环境中评估原生、数组和Flux策略。固定工作负载包含4823个LASTZ任务,本初步实验中每种配置各进行一次干净启动测试。结果显示存在walltime-RPC权衡:在Phoenix上,数组和HyperQueue耗时0.32小时达成目标,Flux耗时0.65小时,但每1000个终端任务可归因RPC降至1396个,三者均为非支配性方案;在Dev环境中,数组和Flux形成观测到的前沿:数组耗时0.80小时,Flux耗时0.84小时,Flux将RPC降至每1000个终端任务1769个。该方法使高性能计算(HPC)站点能够使用同时捕捉用户体验和调度器影响的指标对比部署策略,并可在站点特定RPC预算内选择最快策略。
英文摘要
Wide Nextflow fan-outs on shared Slurm clusters can submit tens of thousands of short tasks. Deployment settings route them through individual jobs, arrays, or nested schedulers inside enclosing allocations. These settings determine workflow turnaround and RPC volume, a shared cost that can degrade scheduler responsiveness. Existing comparisons evaluate whole workflow systems, while per-task queueing metrics cannot span architectures that dispatch inside existing allocations. We contribute a reproducible measurement protocol and benchmark harness. A clean-start clock begins before backend startup or allocation requests, placing architecturally different backends on a common time axis. Per-user Slurm sdiag counters attribute request count as the primary RPC demand measure and controller processing time as sensitivity context, separate from cluster-wide state. We apply the method to Slurm native dispatch, Slurm job arrays, HyperQueue, and Flux on the shared ASU Phoenix production cluster and a single-user Dev cluster. On Phoenix, every aggregation strategy improves both objectives relative to native dispatch; Flux has the lowest RPC demand, while HyperQueue's fastest median walltime is not stable across replicates. On Dev, arrays and Flux improve walltime, while HyperQueue trades the lowest RPC demand for the slowest completion. The method lets an HPC site compare deployment strategies using both user-visible performance and scheduler impact, then select the fastest strategy within its own RPC-demand limit.