编译器到加速器映射的验证:面向机器学习加速器
Verification of Compiler-to-Accelerator Mappings for Machine Learning Accelerators
浏览论文内容
中文总结 AI 辅助
本文提出BOLT框架,首次正式验证ML加速器编译器到加速器映射的正确性,通过循环对齐与数据布局关联,无需编译器额外信息,并在两个开源加速器上成功验证。
中文摘要 AI 辅助
为了满足现代机器学习(ML)应用的性能需求,ML编译器框架支持编译器到加速器的映射,将应用代码的部分内容卸载到专用硬件加速器中的操作上。然而,这些框架大多未能在硬件层面验证这些映射,可能导致功能不匹配。在本文中,我们提出了BOLT,这是第一个用于正式验证ML加速器中粗粒度内联函数的编译器到加速器映射正确性的框架,该验证基于正式的硬件语义。BOLT不需要来自编译器的额外信息,并验证应用代码与映射的硬件加速器内联函数代码之间的功能等价性,包括处理复杂的循环嵌套和硬件中的张量数据布局。它有效地利用了一种模式:先将软件循环与硬件*对齐*,再*关联*相应的数据布局,从而通过良好对齐的乘积程序实现验证。为支持这些步骤,我们提出了两个自定义模板——同步骨架(sync-skeleton)和布局草图(layout-sketch)——分别用于指导用户对齐循环和指定数据布局关系。我们为BOLT开发了一个概念验证原型,并使用它成功验证了两个近期开源ML加速器的多个复杂映射的正确性。
英文摘要
To meet the performance needs of modern machine learning (ML) applications, ML compiler frameworks support compiler-to-accelerator mappings that offload parts of application code to operations in specialized hardware accelerators. However, most of these frameworks do not verify these mappings down to the hardware level, potentially resulting in functional mismatches. In this paper we propose BOLT, the first framework for formally verifying the correctness of compiler-to-accelerator mappings for coarse-grained intrinsics in ML accelerators, with respect to a formal hardware semantics. BOLT does not require additional information from the compiler, and verifies the functional equivalence of the application code and the code for the mapped hardware accelerator intrinsic, including handling of complex loop nests and tensor data layouts in hardware. It effectively utilizes a pattern of *aligning* software loops with the hardware, followed by *relating* corresponding data layouts, to enable verification using well-aligned product programs. To support these steps, we propose two custom templates --- the sync-skeleton and the layout-sketch --- to guide users in aligning loops and specifying data layout relationships, respectively. We have developed a proof-of-concept prototype for BOLT and use it to successfully verify the correctness of several complex mappings for two recent open-source ML accelerators.
发表机构
- Princeton University(普林斯顿大学)
- Meta Platforms(Meta平台)
- Intel INT31(英特尔INT31)
- University of Washington(华盛顿大学)
- Stanford University(斯坦福大学)
- Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。