发表机构
University of Bologna(博洛尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 OpenMP 元降级方法,利用 MLIR 和 DSL 将降级可编程化,在保持性能的同时减少代码量,实现跨目标的高效并行代码生成。
AI 中文摘要
并行硬件的日益多样化对现有编译流程提出了挑战。虽然 OpenMP 为共享内存并行性提供了可移植的抽象,但现有编译器将前端语义与固定的降级策略紧密耦合。这种设计限制了跨不同运行时和架构的性能可移植性。在本文中,我们提出了一种基于多级中间表示(MLIR)框架的模块化并行代码生成方法。我们的方法结合了 OpenMP 前端与一种领域特定语言(DSL),该语言指定了 OpenMP 构造如何针对目标进行降级。我们将这种方法称为 OpenMP 元降级,将降级视为可编程组件而非编译器特定逻辑。这种设计为跨不同目标的代码勾勒、数据共享和运行时接口提供了显式控制。我们在 PolyBench/C-OMP 上针对通用和嵌入式多核目标评估了我们的方法。结果表明,所提出的方法在代码大小开销可忽略(<0.7%)的情况下,与最先进的编译器工具链性能相匹配。将 OpenMP 构造保留到后期降级阶段,使得通用和 OpenMP 特定的 MLIR 优化都能得以实现。通过将降级与编译器内部解耦,我们的方法以 5,182 行代码表达了相同的 OpenMP 子集,比 Clang 中相应的降级代码少约 32%,比 GCC 少约 76%,而支持 pmsis 运行时仅需几行规范,从而无需修改编译器即可快速支持新的运行时。
英文摘要
The increasing diversity of parallel hardware challenges existing compilation flows. While OpenMP provides a portable abstraction for shared-memory parallelism, existing compilers tightly couple the frontend semantics with fixed lowering strategies. This design limits performance portability across different runtimes and architectures. In this paper, we present a modular approach to parallel code generation based on the Multi-Level Intermediate Representation (MLIR) framework. Our approach combines an OpenMP frontend with a domain-specific language (DSL) that specifies how OpenMP constructs are lowered for a target. We call this approach OpenMP meta-lowering, treating lowering as a programmable component rather than compiler-specific logic. This design provides explicit control over code outlining, data sharing, and runtime interfacing across diverse targets. We evaluate our approach on PolyBench/C-OMP across general-purpose and embedded multicore targets. Our results show that the proposed method matches the performance of state-of-the-art compiler toolchains with negligible code-size overhead (< 0.7%). Preserving OpenMP constructs until late lowering stages enables both general-purpose and OpenMP-specific MLIR optimizations. By decoupling lowering from compiler internals, our approach expresses the same OpenMP subset in 5,182 lines of code, ~32% fewer than the corresponding lowering code in Clang and ~76% fewer than in GCC, while supporting the pmsis runtime costs only a few specification lines, enabling rapid support for new runtimes without modifying the compiler.
Comments16 pages, 8 figures, 2 tables. Accepted at CGO 2027 (Salt Lake City, UT, USA)