arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17533cs.SCcs.DS

GPU加速的$\mathbb{F}_2$上快速矩阵乘法搜索

GPU-Accelerated Search for Fast Matrix Multiplication over $\mathbb{F}_2$

Zeke Medley, Anagha Gokul, Quan Luu, Panagiotis Manolios

首次发表
浏览论文内容

中文总结 AI 辅助

提出GPU加速搜索矩阵乘法翻转图算法,每秒超十亿步,证明$\mathbb{F}_2$上翻转图强连通,找到$7\times7$矩阵乘法245次乘法的新纪录。

中文摘要 AI 辅助

我们提出了一种GPU加速算法,用于搜索以少量标量乘法实现矩阵乘法的方式。该算法搜索矩阵乘法翻转图,在NVIDIA H200上每秒执行超过十亿步搜索,相比之前针对$7\times7$矩阵乘法张量的GPU加速搜索,实现了$1000\times$的提升。我们的第二个贡献是证明了在$\mathbb{F}_2$上,有向矩阵乘法翻转图仅通过翻转边和加边即强连通。这消除了先前工作中连通性论证所需的计算昂贵的归约边。利用这种GPU加速搜索程序在我们的简化翻转图上,我们找到了一种在$\mathbb{F}_2$上使用245次乘法实现$7\times 7$矩阵乘法的方法,比之前的记录减少了三次乘法。

英文摘要

We present a GPU-accelerated algorithm for searching for ways to multiply matrices with few scalar multiplications. Our algorithm searches the matrix multiplication flip graph and, on an NVIDIA H200, performs over a billion search steps per second, a $1000\times$ improvement over previous GPU-accelerated search on the tensor for $7\times7$ matrix multiplication. Our second contribution is a proof that, over $\mathbb{F}_2$, the directed matrix multiplication flip graph is strongly connected with only flip and plus edges. This removes the need for computationally expensive reduction edges in the connectivity argument used in previous work. Using this GPU-accelerated search procedure on our simplified flip graph, we find a way to multiply $7\times 7$ matrices over $\mathbb{F}_2$ using 245 multiplications, a three-multiplication improvement over the previous record.

发表机构

  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑