发表机构
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出代价引导框架CREDIT,通过三项创新识别DSMEM有益的工作负载并预测其收益范围,经评估在6种工作负载上较基准实现1.318-1.466倍加速,收益预测准确率达91.7%。
AI 中文摘要
NVIDIA分布式共享内存(DSMEM)支持线程块集群内的直接共享内存访问,但集群同步、远程访问及资源开销使得难以确定DSMEM何时能提升性能。为填补该空白,本文提出CREDIT这一代价引导框架,可识别DSMEM有益的工作负载模式、预测其收益范围,并在各类工作负载上实现一致加速。CREDIT融合三项创新:一是基于剖析的表征方法,识别可能从DSMEM获益的工作负载模式;二是将DSMEM应用于归约-复用工作负载的转换方法;三是基于剖析数据的代价模型,用于确定其收益范围。对各类工作负载的评估显示,CREDIT在收益预测上达到91.7%的准确率;在全部6种工作负载上,CREDIT优于Triton及优化后的非DSMEM CUDA基线,在RTX 5090上的几何平均加速比为1.466倍,在H100上为1.318倍。CREDIT的源代码已公开提供。
英文摘要
NVIDIA distributed shared memory (DSMEM) enables direct shared-memory access within a thread block cluster. However, cluster synchronization, remote access, and resource costs make it difficult to determine when DSMEM improves performance. To fill this gap, we propose CREDIT, a cost-guided framework that identifies DSMEM-profitable workload patterns, predicts their profitability range, and delivers consistent speedups across diverse workloads. CREDIT combines three innovations: (1) a profiling-driven characterization that identifies workload patterns likely to benefit from DSMEM; (2) a transformation that applies DSMEM to reduction-reuse workloads; (3) a cost model based on profiling data, to determine its profitability range. Evaluations on diverse workloads show CREDIT achieves 91.7% prediction accuracy on profitability. CREDIT beats torch.compile, Triton, and optimized non-DSMEM CUDA baselines on all six workloads, with geometric-mean speedups of 1.466x on RTX 5090 and 1.318x on H100. CREDIT's source code is publicly available at https://github.com/zhengxiongli08/CREDIT.
CommentsAccepted by HPEC 2026