发表机构
Illumina(Illumina)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在 Apple M6 芯片上验证 FFT 性能由数据移动而非算术决定,提出寄存器驻留内核,较 MPSGraph 快 2.2 倍,SAR 成像端到端提速 14–19 倍。
AI 中文摘要
快速傅里叶变换(FFT)是雷达、成像和科学计算的基础。快速 GPU FFT 的一个经典规则是计算能容纳在芯片上的最大块,并由这些块组合出更大的变换。我们在 Apple M6 芯片上测试了这一规则,该芯片的 GPU 在从主内存读取一个字节的时间内可执行 29 次算术运算,并且其 GPU 和 CPU 都增加了矩阵单元。将我们的内核与 Apple 的 MPSGraph、vDSP 库以及 MLX 进行比较,我们发现决定 FFT 速度的是数据移动,而非算术运算。在每个 GPU 库中,大批量运行都达到或接近主内存带宽极限,而将数据保留在缓存中的基准测试将实际吞吐量高估了多达 3.7 倍。片上规则仍然能预测速度崩溃的位置:一旦变换超出 GPU 的 32 KiB 本地内存,MPSGraph 和 MLX 的速度就会损失一半或更多。将这些变换保留在寄存器中可避免额外的内存往返:比 MPSGraph 快 2.2 倍,使用半精度存储时快 4.4 倍。GPU 的矩阵单元没有帮助,因为将 FFT 重新表述为矩阵乘积所增加的算术量与其节省的相当。CPU 的矩阵单元比其向量单元快 47 倍,确实有帮助:我们为其编写的内核比 vDSP 快多达 5.3 倍。端到端地,一幅 4096×4096 的合成孔径雷达图像在 GPU 上耗时 8.1 毫秒(使用半精度中间结果时为 6.0 毫秒),比 CPU 参考实现快 14–19 倍。将 CPU 的矩阵单元与 GPU 并行运行并无收益:两者共享同一内存带宽。
英文摘要
The fast Fourier transform (FFT) underlies radar, imaging and scientific computing. A classic rule for fast GPU FFTs is to compute the largest block that fits on chip and compose larger transforms from such blocks. We test this rule on Apple's M6 chip, whose GPU performs 29 arithmetic operations in the time it reads one byte from main memory, and which adds matrix units to both its GPU and CPU. Comparing our kernels with Apple's MPSGraph and vDSP libraries and MLX, we find that data movement, not arithmetic, sets FFT speed. Large batches run at or near the main-memory bandwidth limit in every GPU library, and benchmarks that keep data in cache overstate real throughput by up to $3.7\times$. The on-chip rule still predicts where speed collapses: MPSGraph and MLX lose half or more of their speed once a transform outgrows the GPU's 32\,KiB local memory. Keeping such transforms in registers avoids an extra trip through memory: $2.2\times$ faster than MPSGraph, and $4.4\times$ with half-precision storage. The GPU's matrix units do not help, because recasting the FFT as matrix products adds as much arithmetic as they save. The CPU's matrix unit, 47 times faster than its vector units, does: our kernel for it beats vDSP by up to $5.3\times$. End to end, a $4096\times4096$ synthetic aperture radar image takes 8.1\,ms on the GPU (6.0\,ms with half-precision intermediates), $14$--$19\times$ faster than a CPU reference. Running the CPU's matrix unit alongside the GPU gains nothing: both share one memory bandwidth.