arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08739cs.DC

Tools-CC-Bench:面向HPC与AI工作负载的带压缩的集合通信基准测试套件

Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads

Haozhe Fan, Wei Wang, Xingchen Liu, Man Liu, Xingjian Tian, Haoquan Long, Zedong Liu, Daran Sun, Jinwu Yang, Bo Yang, Jie Liu, Yonggang Che, Hairui Zhao, Guangming Tan, Dingwen Tao

首次发表
浏览论文内容

中文总结 AI 辅助

针对HPC和LLM分布式训练中通信瓶颈,提出CC-Bench基准套件,通过声明式建模、函数级拦截和硬件监控,在CPU/GPU集群上评估三个压缩通信库,揭示精度-性能权衡以指导部署优化。

中文摘要 AI 辅助

分布式HPC和LLM工作负载日益需要高效的通信以实现可扩展性,然而不断增长的数据移动已成为主要的性能瓶颈。通信压缩可以减少这种开销并补充执行级优化,但其益处仍难以评估,因为现有基准测试缺乏对多样化后端、真实数据集、应用特定精度指标以及重叠引发的资源争用的支持。我们提出CC-Bench,一个轻量级、可扩展且面向应用的基准测试套件,用于在真实执行条件下评估通信压缩。CC-Bench采用声明式应用环境建模,将性能分析逻辑与通信库、数据集和保真度指标解耦,从而实现跨库的可移植评估。它进一步结合函数级拦截和硬件计数器监控,以刻画每阶段延迟、硬件利用率、数值保真度和计算干扰。借助来自HPC和LLM工作负载的代表性数据集,CC-Bench在CPU和GPU集群上评估了三个支持压缩的通信库,揭示了精度-性能权衡和瓶颈,以指导实际部署和优化。

英文摘要

Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. Communication compression can reduce this overhead and complement execution-level optimizations, but its benefits remain difficult to assess because existing benchmarks lack support for diverse backends, realistic datasets, application-specific accuracy metrics, and overlap-induced resource contention. We present CC-Bench, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions. CC-Bench uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation. It further combines function-level interception and hardware counter monitoring to characterize per-phase latency, hardware utilization, numerical fidelity, and computation interference. With representative datasets from HPC and LLM workloads, CC-Bench evaluates three compression-enabled communication libraries on CPU and GPU clusters, revealing accuracy-performance trade-offs and bottlenecks to guide practical deployment and optimization.

发表机构

  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)
  • School of Computer Science, Nanjing University(南京大学计算机系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑