arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在真实硬件上验证EPYC猜想:在NCSA Delta(AMD EPYC 7763 Milan)上进行MoA引导的稠密矩阵乘法

Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta (AMD EPYC 7763 Milan)

Lenore M Mullin

arXiv 2608.11533首次发表:更新:

AI 中文总结

本文在NCSA Delta硬件上验证EPYC猜想,提出将共驻等视为形状参数,得出MoA引导的稠密矩阵乘法块大小M_C=256性能优于传统值,且MoA流水线在NUMA环境下吞吐量损失更小的结论。

AI 中文摘要

一项配套实证研究曾推测,MoA引导的稠密矩阵乘法需要在AMD EPYC Bergamo上按每个CCD进行重新校准;由于无法获得该机器的使用权限,本文报告了在架构相关的芯片NCSA Delta上进行的测试。本文的核心主张:此处的所有结果均源于形状,而非抽象意义上的架构——缓存形状、访问模式形状,以及共驻(co-tenancy)本身被视为一种形状参数。修正后的块大小M_C=256直接源自Delta的实际512 KB L2缓存除以数据类型自身的字节宽度,无需事后拟合任何参数,且其性能比从M1 Pro继承而来的M_C=64高出30%至59%——块的形状终于匹配了其原本应占据的缓存形状。同样的推理可扩展至共享层面:将共驻视为原始容量会产生一个数量级的误差,而将共驻视为形状参数(容量除以并发消费者数量)则能得出合理结果。一项受控NUMA实验分离出了第三种形状:GEBP与MoA流水线计算出的结果完全相同,且达到的带宽也一致,但GEBP因远程内存损失了13.9%的吞吐量,而MoA流水线仅损失2.4%——这并非速度差异,而是每个内核的访问模式在内存中所呈现的形状不同。核心结果再现了M1 Pro在四场测试中赢了三场的表现,而AWS Graviton4未实现这一模式。有一个问题被提出但未解决:Delta上MoA相对于GEBP的更大优势是源于真正的速度提升,还是源于其利用率占比更高。此处的所有结果均是在测量前预测的,而非事后拟合的。

英文摘要

A companion empirical study conjectured that MoA-guided dense matrix multiplication would need per-CCD recalibration on AMD EPYC Bergamo; allocation access to that machine was declined, and this paper reports the resulting test on NCSA Delta, an architecturally related chip. Its central claim: every result here traces to shape, not architecture in the abstract -- cache shape, access-pattern shape, and co-tenancy itself treated as a shape parameter. The corrected block size M_C = 256 follows directly from Delta's real 512 KB L2 divided by the data type's own byte width, no parameter fit after the fact, and outperforms the M1-Pro-inherited M_C = 64 by 30--59% -- a block's shape finally matching the cache shape it was always meant to occupy. The same reasoning extends to a shared level: co-tenancy treated as raw capacity fails by an order of magnitude, while co-tenancy treated as a shape parameter -- capacity divided by concurrent consumers -- survives. A controlled NUMA experiment isolates a third shape: GEBP and MoA-pipelined compute the identical result at identical achieved bandwidth, yet GEBP loses 13.9% of its throughput to remote memory while MoA-pipelined loses only 2.4% -- not a speed difference, a difference in the shape each kernel's access pattern traces through memory. The headline result reproduces the M1 Pro's exact three-of-four win over a Strassen-GEBP hybrid, a pattern AWS Graviton4 did not achieve. One question is named rather than resolved: whether Delta's larger MoA-over-GEBP margin reflects genuine speed or an unmatched utilization fraction. Every result here was predicted before it was measured, not fitted after.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑