arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05138cs.LGcs.DC

从80倍到385倍:在对称调优基准下,基于L2带宽上限的最佳匹配单元搜索

From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

Andrew James Amos

首次发表
浏览论文内容

中文总结 AI 辅助

本文在对称调优cuSPARSE基准的条件下,调优SparseBin算法的最佳匹配单元搜索,实现了与早期CUDA实现性能差距从80倍到385倍的提升,同时揭示了L2带宽上限的相关特性。

中文摘要 AI 辅助

GPU实现的对比通常是不对称的:一方由其作者调优,另一方则直接使用默认设置运行。本文报告了一项调优工作,同时调优了一种新型SOM算法(SparseBin)及其对比基准算法cuSPARSE。自组织映射训练中占主导地位的最佳匹配单元搜索,通过四个调优手段——块大小、块成员聚类、神经元轴分块和向量化加载——实现了每轮训练较之前已发表配置5.6至10.1倍的加速,映射尺寸从32×32到512×512不等;同时,其与早期MEDLINE图集背后CUDA实现的性能差距从约80倍提升至约385倍。与SparseBin对比的cuSPARSE,其每一项调优手段均有对应的类似操作,最终实现了2至3倍的加速。调优后的内核在峰值的77%处达到L2带宽上限,其余单元则在40%至65%之间,这一结果是终端结果而非中间节点;所有未测试的调优手段要么受限于该上限,要么经测量无效果。

英文摘要

Comparisons between GPU implementations are usually asymmetric: one side is tuned by its author, the other is run as found. I report a programme that tuned both a novel SOM algorithm (SparseBin) and the baseline algorithm it was being compared to (cuSPARSE). The best-matching-unit search that dominates self-organizing map training was tuned through four levers - tile size, tile-membership clustering, neuron-axis chunking and vectorised loads - reaching 5.6-10.1x per epoch over the previously published configuration at map sizes from 32x32 to 512x512, and lifting the margin over the CUDA implementation behind our earlier MEDLINE atlases from ~80x to ~385x. cuSPARSE, the implementation SparseBin is compared against, received every lever with an analogue on its side, and became 2-3x faster in the process. The tuned kernel pressed the L2 bandwidth roof at 77% of peak with every other unit at 40-65%, bounding any further lever at ~1.3x - a terminal result rather than a waypoint, and every untested lever was either capped by that bound by construction or measured null.

发表机构

  • College of Medicine and Dentistry, James Cook University(詹姆斯库克大学牙医学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑