基于Flash Attention的快速高斯和计算
Fast Gauss Sums via Flash Attention
- Faculty of Mathematics(数学学院)
- Chemnitz University of Technology(开姆尼茨工业大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出通过两次输入增强,利用Flash Attention计算带任意符号权重的高斯核和,无需自定义GPU代码,在fp16、D>8时性能优于PyTorch与PyKeOps,内存扩展呈线性。
中文摘要 AI 辅助
高斯核和是最大均值差异(MMD)、核梯度流、斯坦变分梯度下降(SVGD)等众多核方法的计算核心。同时,softmax注意力经过大量硬件感知代码工程优化,最终形成了Flash Attention。我们证明,带有任意符号权重的高斯核和可通过Flash Attention计算:仅需两次小型输入增强,就能将归一化的softmax归约转换为非归一化的高斯和,无需编写一行自定义GPU代码。对于fp16精度下特征维度D>8的情况,该方法在速度、内存开销和精度上均优于编译后的PyTorch代码及PyKeOps内核(通常优势显著),且其内存扩展保持线性。
英文摘要
Gaussian kernel sums are the computational core of maximum mean discrepancies (MMDs), kernel gradient flows, Stein variational gradient descent (SVGD), and many other kernel methods. At the same time, softmax attention has received an extraordinary amount of hardware-aware code engineering, culminating in flash attention. We show that Gauss kernel sums with arbitrary, signed weights can be evaluated via flash attention: two small input augmentations turn the normalized softmax reduction into the unnormalized Gauss sum, without writing a single line of custom GPU code. For feature dimension D>8 in fp16, this approach beats compiled PyTorch code as well as PyKeOps kernels (often significantly) in speed, memory-overhead and accuracy. Indeed, its memory scaling remains linear.