arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30099cs.AR

编译器与硬件协同设计:面向加速器架构

Compiler and Hardware Co-Design for Accelerator Architectures

  • Technical University of Denmark(丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Karl Herman Krause, Emad Jacob Maroun, Martin Schoeberl

AI总结:

针对异构加速器全栈集成难的问题,提出EAAC编译器与硬件协同设计架构,基于MLIR支持快速原型,原型GEMM加速器较RISC-V基线实现28倍加速,并指出信号量分配优化方向。

AI中文摘要:

异构加速器架构为计算密集型工作负载提供了一条高效的性能路径。然而,全栈集成仍然困难。我们提出了EAAC(可扩展加速器架构),这是一种灵活且可扩展的编译器与硬件架构,旨在降低硬件-编译器协同开发的开销,以支持硬件加速器的快速原型设计。EAAC面向具有可预测内存访问模式的静态数据流工作负载。通过使用MLIR,我们实现了与多种可生成MLIR的前端的潜在集成。通过提供一组最小的编译器功能来实现必要的数据编排,我们降低了让一个简单的硬件加速单元实现并运行所需的努力,同时为后续工作提供了一个相对空白且无预设观点的起点。我们通过一个原型GEMM加速器验证了该方法,该加速器将脉动阵列单元与嵌入式RISC-V内核相结合,并通过EAAC MLIR流水线进行端到端编译。在一个合成的全连接层工作负载上,该加速器相比仅使用RISC-V的基线实现了28倍的执行时间加速,其中GEMM操作本身仅占总执行时间的1.64%。我们进一步表征了编译器的硬件信号量分配,表明在最坏情况下所需信号量数量与指令数量呈线性增长,并将此确定为未来优化的具体目标。

英文摘要:

Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardware-compiler co-development for rapid prototyping of hardware accelerators. EAAC targets static data-flow workloads with predictable memory access patterns. By using MLIR, we enable possible integration with a range of different frontends that emit MLIR. And by providing a minimal set of compiler functionality that enable necessary data-orchestration we lower the effort needed to get a simple implementation of a hardware acceleration unit up and running, while providing a relatively blank and un-opinionated starting point for further work. We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core, compiled end-to-end through the EAAC MLIR pipeline. On a synthetic fully-connected-layer workload, the accelerator achieves a 28x speedup in execution time over a RISC-V-only baseline, with the GEMM operation itself accounting for only 1.64\% of total execution time. We further characterize the compiler's hardware-semaphore allocation, showing that the number of semaphores required scales linearly with instruction count in the worst case, and identify this as a concrete target for future optimization.

补充信息

↑