arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SparseDitto:基于大语言模型的智能体系统,为不同稀疏性模式定制GPU内核

SparseDitto: An Agentic Sparse Compilation Framework through Architecture-Aware Synthesis on GPUs

Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding

arXiv 2608.05033首次发表:更新:

发表机构

University of Minnesota; Northeastern University(明尼苏达大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SparseDitto是基于LLM的系统,可为不同矩阵、算子和GPU定制稀疏矩阵GPU内核,在多类算子与GPU上较cuSPARSE实现显著加速,还可提升GCN训练速度。

AI 中文摘要

稀疏矩阵内核是科学计算、图分析和机器学习的基础,其GPU性能高度依赖输入稀疏性模式与执行策略。对于同一矩阵上的相同SpMM操作,cuSPARSE在CSR格式与Blocked-ELL格式间存在350倍的性能差距。本研究对多种数据格式、专用系统和稀疏编译器的分析表明,不存在单一实现能在所有稀疏性模式和算子上始终占据优势,这促使我们开发一种可针对每个工作负载和目标GPU调整其表示、执行策略与硬件映射的系统。我们提出SparseDitto,这是一个基于大语言模型的系统,可为每个矩阵、算子和目标GPU构建对应的GPU内核,在统一设计框架内支持SpMV、SpMM和SpGEMM三类算子。该系统先通过轻量级加法模型,利用输入矩阵的结构特征对已有的策略进行排序;再由感知架构的规划器提出若干候选设计;编码与验证智能体则借助目标GPU的实测数据实现并优化这些设计。在三类稀疏算子和多样矩阵组成的测试集上,SparseDitto在NVIDIA RTX PRO 6000 GPU上较cuSPARSE实现了2.68倍的几何平均加速比,最高达146.61倍;在NVIDIA H200 GPU上则实现2.79倍的几何平均加速比,最高达78.5倍。其生成的SpMM内核还可将全批量GCN训练加速最高3.39倍。

英文摘要

Sparse matrix computation performance on GPU depends on how representation and execution schedule match the input structure and target hardware. No single implementation consistently dominates across sparsity patterns, operators, and hardwares. Existing sparse compilers and specialized systems cannot cover all of them simultaneously. We present SparseDitto, an agentic sparse compilation framework for sparse matrix computation on GPUs. It jointly synthesizes representation, execution schedule, and hardware mapping in a unified compilation plan. Structural analysis and a learned template-ranking prior guide architecture-aware synthesis. LLM-guided lowering realizes each plan as CUDA code, while target-GPU profiling drives plan refinement. SparseDitto covers multiple operators, e.g., SpMV, SpMM, and SpGEMM, and various representations within one framework. It can also automatically adapt to different hardwares. Across various SuiteSparse matrices, SparseDitto achieves geometric-mean speedups over cuSPARSE of $2.68\times$ on an NVIDIA RTX PRO 6000 and $2.79\times$ on an NVIDIA H200 (up to 146.61$\times$). Its generated SpMM kernels accelerate full-batch GCN training by up to $3.39\times$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑