发表机构
University of Illinois Urbana-Champaign; National Center for Supercomputing Applications, University of Illinois Urbana-Champaign; University of Southern California; California Institute of Technology(伊利诺伊大学厄巴纳-香槟分校; 伊利诺伊大学厄巴纳-香槟分校国家超级计算应用中心; 南加州大学; 加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GEM-KMeans提出谱归一化等价NLR公式,融合更新与投影于矩阵乘法尾声,仅需单因子存储,实现海量数据内存高效且精确的GPU聚类。
AI 中文摘要
在不牺牲统计精度的情况下,对聚类问题进行内存高效扩展是大规模数据分析和机器学习问题的核心关注点。用于K均值聚类的非负低秩(NLR)矩阵分解是一种可扩展的聚类方法,它与具有最优平均情况精确恢复保证的半定松弛相关联。然而,直接的GPU实现NLR需要多个大型因子大小的缓冲区和大量本质上受内存限制的数据移动。在本文中,我们引入了GEM-KMeans,这是一种谱归一化但数学上等价的NLR公式,它将梯度更新、非负投影以及用于归一化和迭代移动的充分统计量融合到矩阵乘法尾声中。与保留三个大型因子大小的数组不同,我们的IO感知GPU实现仅物化一个单一因子,并在高带宽内存(HBM)中使用小型瓦片缩减数组作为额外存储。我们推导了用于优化聚类目标函数的显式内存成本和谱归一化平滑度界限。在合成和真实数据集上展示了大规模下的精确聚类,其中GEM-KMeans相对于现有GPU加速的Lloyd算法的性能提升涉及数据相关的运行时权衡。
英文摘要
Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for $K$-means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue. Instead of retaining three massive factor-sized arrays, our IO-aware GPU implementation materializes only one single factor with small tile-reduction arrays as additional storage in the High Bandwidth Memory (HBM). We derive explicit memory costs and spectrally normalized smoothness bounds for optimizing the clustering objective function. Accurate clustering is demonstrated at massive scales on synthetic and real datasets, where performance gains of GEM-KMeans over existing GPU-accelerated Lloyd's algorithms involve data-dependent runtime tradeoffs.