AI 中文总结
针对空间数据流架构难以处理非结构化网格计算的问题,提出联合代码与数据分解方法,通过通信和内存建模,自动化分析代码,定义高维分解并应用内存优化技术,将LULESH应用映射到相关引擎,证明该方法能让非结构化网格代码超越GPU。
AI 中文摘要
空间数据流架构是高性能计算中新兴的硬件模式,其网格连接的固定内存处理元素专为具有二维邻域的结构化网格内核量身定制。然而,实际的多物理场代码通常在非结构化网格上计算,这会导致间接内存访问和高维通信模式,使其无法直接映射到上述架构。本文采用以模型为中心的原则性方法,将非结构化问题划分到空间数据流架构上。通过通信和内存建模,我们提出了一种联合分解方法,该方法同时考虑应用程序的字段大小及其子例程。特别是,我们自动化了原始代码的分析过程,定义了一种通过空间填充曲线最小化通信的高维分解方法,并应用了在这种内存受限环境中至关重要的内存优化技术。我们展示了将利弗莫尔非结构化拉格朗日显式激波流体动力学(LULESH)应用程序映射到Cerebras晶圆级引擎上,表明更大的非结构化网格代码仍然可以超越GPU。
英文摘要
Spatial Dataflow Architectures are an emerging hardware pattern in high-performance computing, whose mesh-connected fixed-memory processing elements are tailored for structured grid kernels with two-dimensional neighborhoods. However, practical multiphysics codes are often computed on unstructured grids, which induce indirect memory accesses and high-dimensional communication patterns, making them infeasible to directly map onto said architectures. This work takes a principled, model-centric approach to partitioning unstructured problems onto spatial dataflow architectures. Through communication and memory modeling, we propose a joint decomposition that considers both the size of the application's fields and its subroutines. In particular, we automate the analysis process of the original code, define a high-dimensional decomposition that minimizes communication via space-filling curves, and apply memory optimization techniques, crucial in this memory-limited environment. We demonstrate mapping the Livermore Unstructured Lagrangian Explicit Shock Hydrodynamics (LULESH) application to the Cerebras Wafer-Scale Engine, showing that larger, unstructured grid codes can still outperform GPUs.
Comments11 pages, 12 figures