arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于FPGA加速器设计的模式引导设计空间探索

Pattern-Guided Design Space Exploration for FPGA Accelerator Design

Jialiang Zhang, Weiman Yan, Yuelin Zou

arXiv 2607.15068首次发表:更新:

AI 中文总结

研究针对FPGA加速器设计中因调度决策产生的组合设计空间问题,提出PATTERNDSE框架,将重复计算模式映射到调度空间,经多步骤评估候选调度,在六个内核上验证其能大幅减少候选者数量且保持最佳延迟,提升设计效率。

AI 中文摘要

高级综合(HLS)将FPGA加速器设计的抽象级别从硬件描述语言提升到C/C++,但其高质量结果仍依赖于诸如流水线、展开、平铺、重排序和缓冲等调度决策。这些决策产生了组合设计空间,而许多数值内核呈现出重复的计算模式,暗示了不同的优化策略。本文提出了PATTERNDSE,一个轻量级的模式引导设计空间探索(DSE)框架,用于在Allo(一个面向调度的HLS编程系统)中编写的FPGA内核。PATTERNDSE将重复的计算模式映射到紧凑的调度空间,应用候选调度,通过LLVM执行验证功能正确性,检查HLS C代码生成,并在Vitis HLS合成之前使用简单的模式感知估计器对候选者进行排名。我们在六个代表性内核上评估了PATTERNDSE。与穷举精简基线相比,模式引导的DSE将HLS评估的候选者数量从140减少到29,总体搜索减少了4.83倍,单个内核最多减少了12.0倍。在所有评估内核中,PATTERNDSE恢复了与穷举精简基线相同的最佳有效Vitis HLS延迟,表明计算模式信息可以在保留高质量HLS结果的同时修剪无生产性的调度组合。

英文摘要

High-level synthesis (HLS) raises the abstraction level of FPGA accelerator design from hardware description languages to C/C++, but high-quality results still depend on schedule decisions such as pipelining, unrolling, tiling, reordering, and buffering. These decisions create a combinatorial design space, while many numerical kernels exhibit recurring computation patterns that suggest different optimization strategies. This paper presents PATTERNDSE, a lightweight pattern-guided design space exploration (DSE) framework for FPGA kernels written in Allo, a scheduling-oriented HLS programming system. PATTERNDSE maps recurring computation patterns, including elementwise maps, reductions, matrix-vector operations, matrix-matrix operations, and stencil-like updates, to compact schedule spaces. It then applies candidate schedules, validates functional correctness through LLVM execution, checks HLS C code generation, and uses a simple pattern-aware estimator to rank candidates before Vitis HLS synthesis. We evaluate PATTERNDSE on six representative kernels: vecadd, axpy, dot, matvec, gemm, and jacobi2d. Compared with an exhaustive-lite baseline, pattern-guided DSE reduces the number of HLS-evaluated candidates from 140 to 29, achieving a 4.83x overall search reduction and up to 12.0x reduction for individual kernels. Across all evaluated kernels, PATTERNDSE recovers the same best valid Vitis HLS latency as the exhaustive-lite baseline, demonstrating that computation-pattern information can prune unproductive schedule combinations while preserving high-quality HLS outcomes.

Comments6 pages, 4 figures, IEEE ICECCME conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑