AI 中文总结
本文指出OpenMP Target Offload的半自动缓冲区同步功能存在正确性、性能等缺陷,说明OpenMP原生支持的缓冲区局部性显式处理可避免多数此类问题,但也有自身不足。
AI 中文摘要
OpenMP Target Offload是一种流行的GPU技术,用于移植处理大型缓冲区的计算代码。通常强调的主要易用性功能之一是分区内存设置中缓冲区同步的半自动处理,这是离散GPU系统的典型特征。然而,该功能存在潜在的正确性、性能和资源消耗缺陷。本文概述了该方法的缺陷,并阐述了通过OpenMP原生支持的缓冲区局部性显式处理如何避免大部分此类陷阱,但同时也存在其自身的缺点。
英文摘要
OpenMP Target Offload is a popular GPU technology for porting compute codes that operate on large buffers. One of the main usability features that is typically emphasized is the semi-automatic handling of buffer synchronization in partitioned memory setups, typical of discreet GPU systems. That feature however comes with potential correctness, performance and resource consumption drawbacks. This paper outlines the drawbacks of that approach and outlines how explicit handling of buffer locality, natively supported through OpenMP, avoids most of those pitfalls, but also comes with its own downsides.
Comments4 pages