arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2408.01391cs.DCcs.LG

FT K-means:一种具有容错能力的GPU高性能K-means算法

FT K-means: A High-Performance K-means on GPU with Fault Tolerance

  • University of California, Riverside(加利福尼亚大学河滨分校)
  • University of Houston(休斯顿大学)
  • Argonne National Laboratory(阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shixun Wu, Yitong Ding, Yujia Zhai, Jinyang Liu, Jiajun Huang, Zizhe Jian, Huangliang Dai, Sheng Di, Bryan M. Wong, Zizhong Chen, Franck Cappello

更新

AI总结:

FT K-means通过逐步优化、模板化代码生成和warp级张量核纠错方案,在GPU上实现了比cuML快10%-300%的高性能K-means,并以仅11%的开销提供在线容错能力。

AI中文摘要:

K-means是一种广泛使用的聚类算法,然而其效率主要受限于距离计算的计算成本。现有实现在计算单元的利用率上存在不足,并且缺乏对软错误的抵御能力。为应对这些挑战,我们提出了FT K-means,一种具有在线容错能力的GPU加速高性能K-means实现。我们首先提出一种逐步优化策略,其性能可与NVIDIA的cuML库相媲美。我们进一步通过基于模板的代码生成框架改进FT K-means,该框架支持不同数据类型并适应不同输入形状。我们提出了一种新颖的warp级张量核纠错方案,以解决现有容错方法在复制操作期间因内存异步而失效的问题。我们在NVIDIA T4 GPU和A100 GPU上的实验评估表明,不带容错功能的FT K-means性能优于cuML的K-means实现,在涉及不规则数据形状的场景中性能提升10%-300%。此外,FT K-means的容错功能仅引入11%的开销,即使在每秒注入数十个错误的情况下仍能保持稳健的性能。

英文摘要:

K-means is a widely used algorithm in clustering, however, its efficiency is primarily constrained by the computational cost of distance computing. Existing implementations suffer from suboptimal utilization of computational units and lack resilience against soft errors. To address these challenges, we introduce FT K-means, a high-performance GPU-accelerated implementation of K-means with online fault tolerance. We first present a stepwise optimization strategy that achieves competitive performance compared to NVIDIA's cuML library. We further improve FT K-means with a template-based code generation framework that supports different data types and adapts to different input shapes. A novel warp-level tensor-core error correction scheme is proposed to address the failure of existing fault tolerance methods due to memory asynchronization during copy operations. Our experimental evaluations on NVIDIA T4 GPU and A100 GPU demonstrate that FT K-means without fault tolerance outperforms cuML's K-means implementation, showing a performance increase of 10\%-300\% in scenarios involving irregular data shapes. Moreover, the fault tolerance feature of FT K-means introduces only an overhead of 11\%, maintaining robust performance even with tens of errors injected per second.

↑