发表机构
University of Leeds; Argonne National Laboratory(利兹大学; 阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对AMD GPU矩阵乘法器不符合IEEE 754标准的问题,通过设计测试向量和MATLAB模型,实现了对CDNA 1/2/3架构的位级精确建模,并量化了与NVIDIA张量核心的精度差异。
AI 中文摘要
近期GPU上可用的矩阵乘法器不符合IEEE 754浮点标准。矩阵乘法器的特性因厂商以及同一厂商的不同架构而异,例如累加器宽度、舍入行为、归一化点、中间下溢和上溢逻辑、次正规数的处理以及特殊输入的处理。因此,小矩阵乘法结果在不同设备间的可复现性无法实现,且无法通过软件控制达成。矩阵乘法器的实现细节未被文档化,使得难以解释计算结果中的差异。我们刻画了三种AMD GPU架构(CDNA 1、CDNA 2和CDNA 3)上矩阵乘法器的数值行为,分别使用MI100、MI210/250和MI300A/300X GPU。我们设计了测试向量以针对所有支持的输入格式的数值特性,并提供了每个向量为何能基于设备输出确定特定数值特性的推导和推理。随后,为每种架构开发了基于MATLAB的矩阵乘法器软件模型,并使用包含1000万组随机输入向量的随机化测试套件,针对硬件验证了位级可复现性。为实现此目标,我们应用了先前开发的技术,在循环中迭代细化模型精度,通过随机测试随后进行测试细化,直到模型在每个测试用例上与硬件匹配。最后,作为模型可用于实验研究的概念验证,我们在两个示范性数值应用中利用了这些模型,量化了AMD矩阵核心与NVIDIA张量核心之间的应用级精度差异。
英文摘要
Matrix multipliers available on recent GPUs do not conform with the IEEE 754 floating point standard. Features of matrix multipliers differ across vendors and architectures of the same vendor, such as accumulator width, rounding behaviour, normalisation points, intermediate underflow and overflow logic, the handling of subnormals, and the treatment of special inputs. As a result, reproducibility of small matrix multiplier results across devices is not possible and cannot be achieved by software control. Implementation details of matrix multipliers are not documented, making it difficult to interpret discrepancies in the computed results. We characterise the numerical behaviour of matrix multipliers across three AMD GPU architectures: CDNA 1, CDNA 2, and CDNA 3, using the MI100, MI210/250, and MI300A/300X GPUs, respectively. We design test vectors to target numerical features for all supported input formats and provide the derivation and the reasoning for why each vector allows to determine a particular numerical feature based on the outputs of the devices. MATLAB-based software models of the matrix multipliers are then developed for each architecture and validated for bit-level reproducibility against hardware using a randomized test suite consisting of 10 million sets of random input vectors. To achieve this, we applied a previously developed technique to iteratively refine the accuracy of the models in a loop, by randomized testing followed by test-refinement until the model matches the hardware for every test case. Finally, as a proof of concept for what experimental research can be done with the models, we have utilised them in two demonstrative numerical applications, quantifying application-level accuracy differences between AMD matrix cores and the NVIDIA tensor cores.