WLCG 迷你能力挑战:主机调优以改进 WAN 数据传输
WLCG Mini-Capability Challenge: Host Tuning to Im- prove WAN Data Transfers
另 6 家 · 查看机构详情
- Vanderbilt University(范德堡大学)
- Brookhaven National Laboratory(布鲁克海文国家实验室)
- University of California, San Diego(加州大学圣迭戈分校)
- University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)
- Energy Sciences Network(能源科学网络)
- Physics Department, University of Michigan(密歇根大学物理系)
- University of Nebraska, Lincoln(内布拉斯加大学林肯分校)
- Michigan State University(密歇根州立大学)
- Fermi National Accelerator Laboratory(费米国家加速器实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对WLCG数据传输瓶颈,通过主机调优(TCP、队列、卸载等)在多个站点实验,发现收益取决于本地瓶颈,并诊断出配置问题,为站点优化提供指导。
中文摘要 AI 辅助
高吞吐量数据移动是 WLCG 计算模型的一项决定性要求,然而端到端性能往往更多地受限于主机级配置,而非网络骨干。我们展示了在 ATLAS 和 CMS 站点(FNAL、UCSD、UNL、BNL、AGLT2、MWT2 和 Vanderbilt)开展的 WLCG 迷你能力挑战的结果,该挑战聚焦于主机优化,使用配备 25 Gbps 或更高网卡的 EL8/9 主机。NET2 参与了随后的跨大西洋测试。该方法结合了 ESnet Fasterdata 指南与 OSG 托管的 this http URL 框架,应用 TCP、队列、卸载和环形缓冲区调优,并支持保存和恢复功能,以实现安全的基线比较。结果表明,调优收益关键取决于本地瓶颈。在 DTN 和存储均具备能力的情况下(UNL),FTS 吞吐量从 60 Gbps 提升至 90 Gbps。在存储为约束条件的情况下(UCSD),iperf3 吞吐量从 20 Gbps 提升至 90 Gbps,而 FTS 保持不变。在网卡为约束条件的情况下(Vanderbilt,25 G 网卡),调优配置反而降低了吞吐量,这促使对每个站点的适用性进行审查。该挑战还发现并解决了 AGLT2 的 dCache 代理配置错误,表明迷你挑战除了原始性能测量外,还承担诊断功能。
英文摘要
High-throughput data movement is a defining requirement of the WLCG computing model, yet end-to-end performance is often constrained less by the network backbone than by host-level configuration. We present the results of a WLCG mini-capability challenge focused on host optimization at ATLAS and CMS sites (FNAL, UCSD, UNL, BNL, AGLT2, MWT2, and Vanderbilt), using EL8/9 hosts with 25\,Gbps or higher NICs. NET2 participated in a subsequent trans-Atlantic test. The methodology combines ESnet Fasterdata guidance with the OSG-hosted \texttt{fasterdata-tuning.sh} framework, applying TCP, queue, offload, and ring-buffer tuning with save and restore support for safe baseline comparisons. Results show that tuning benefit depends critically on the local bottleneck. Where DTNs and storage were both capable (UNL), FTS throughput rose from 60 to 90\,Gbps. Where storage was the binding constraint (UCSD), iperf3 throughput improved from 20 to 90\,Gbps while FTS was unchanged. Where the NIC was the binding constraint (Vanderbilt, 25\,G NIC), the tuning profile actively reduced throughput, motivating a review of per-site applicability. The challenge also surfaced and resolved a dCache proxying misconfiguration at AGLT2, demonstrating that mini-challenges serve a diagnostic function beyond raw performance measurement.