OpenMP中的统一共享内存(USM):在英特尔加速器上的实现、可编程性与性能
Unified Shared Memory in OpenMP: Implementation, Programmability, and Performance on Intel Accelerators
浏览论文内容
中文总结 AI 辅助
本文介绍英特尔实现的OpenMP USM特性,评估其在HPC应用中的采用复杂度与Battlemage GPU上的性能,指出该特性虽有低开销,但部分应用可从中受益,利于快速原型开发与移植。
中文摘要 AI 辅助
OpenMP 5.0通过requires指令引入了统一共享内存(Unified Shared Memory,USM)特性。该特性在加速器与主机之间提供了唯一且通用的地址空间,允许在不同设备上访问(解引用)相同的内存地址,从而避免了为维护地址空间一致性而进行显式数据传输的负担,简化了OpenMP编程模型的采用,便于快速原型开发及将应用移植到带加速器的OpenMP环境中。本文介绍了英特尔对USM的实现,简要讨论其在软件栈(操作系统内核、编译器及运行时)中的实现方式,评估现有使用OpenMP加速器的高性能计算(HPC)应用采用该特性的复杂度,最后评估这些应用在英特尔Battlemage GPU上采用USM时的性能。USM预计不会为已采用优化的显式细粒度数据移动控制的应用带来性能提升,本文结果显示其几何平均开销低于1.2倍(进一步优化后可达到1.03倍),但仍存在从该特性中受益的应用,使其对已移植的应用也具有吸引力。
英文摘要
OpenMP 5.0 introduced the Unified Shared Memory (USM) feature through the requires directive. The feature simplifies the adoption of the OpenMP programming model by providing a unique and common address space between the accelerators and the host and allowing the access (dereference) of the same memory address on different devices, thus avoiding the burden of explicit data transfers to maintain the consistency between the address spaces. Hence, the feature eases quick prototyping and porting of applications to OpenMP with accelerators. In this paper, we introduce the Intel implementation for USM. We briefly discuss its implementation in the software stack (OS kernel, compiler, and runtime), then assess its adoption complexity in existing HPC applications using OpenMP for accelerators, and, finally, evaluate the performance of these applications when adopting USM on an Intel Battlemage GPU. USM is not expected to grant performance uplifts to already optimized applications with explicit, granular data-motion control and our results show an overhead with a geometric mean below 1.2x (1.03x seems achievable with further optimizations). Yet, in this paper we show there exist applications that benefit from this feature, making it attractive even for already ported applications.