发表机构
University of Bologna; CINECA - HPC Department(博洛尼亚大学; CINECA高性能计算部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文利用RISC-V向量扩展优化OpenFOAM中的稀疏矩阵-向量乘法,集成自定义平滑器,在EPAC和Sophon SG2044上分别实现6倍和1.5倍加速。
AI 中文摘要
计算流体动力学(CFD)在很大程度上依赖于基于稀疏线性代数内核的线性求解器的效率。广泛使用的框架如OpenFOAM主要通过MPI在域分解层面利用并行性。对向量/SIMD架构的支持仅限于编译器自动向量化。此外,对此类架构的支持受到OpenFOAM内部矩阵数据格式的限制,该格式本质上不适合向量处理器高效执行所需的连续内存访问。在本工作中,我们聚焦于两种截然不同的RISC-V架构:原型长向量EPAC加速器和商用短向量CPU Sophon SG2044。在这些平台上,我们使用RISC-V向量内联函数优化了稀疏矩阵-向量乘法(SpMV),并将其集成到自定义平滑器中,在运行时将内部数据转换为向量友好的格式。在EPAC测试芯片上的实验结果显示平滑器加速6倍;在配备Sophon SG2044处理器的Monte Cimone(MCv2)集群上的基准测试实现了1.5倍的平滑器加速,证明了传统CFD代码可以在研究性和商业新兴硬件上得到有效加速。
英文摘要
Computational Fluid Dynamics (CFD) relies heavily on the efficiency of linear solvers based on sparse linear algebra kernels. Widely used frameworks like OpenFOAM exploit parallelism primarily at the domain decomposition level via MPI. Support for vector/SIMD architectures is limited to compiler auto-vectorization. Furthermore, support for such architectures is limited by OpenFOAM's internal matrix data format, which is intrinsically ill-suited for the contiguous memory accesses required for efficient execution on vector processors. In this work, we focused on two very different RISC-V architectures: the prototype long-vector EPAC accelerator and the commercial short-vector CPU Sophon SG2044. On these platforms, we optimized the Sparse Matrix-Vector multiplication (SpMV) using RISC-V vector intrinsics and integrated it into a custom smoother, performing a runtime conversion of internal data into a vector-friendly format. Experimental results on the EPAC test chip show a 6x speedup for the smoother; benchmarks on Monte Cimone (MCv2) cluster with the Sophon SG2044 processor achieve a 1.5x smoother speedup, proving that legacy CFD codes can be effectively accelerated on both research and commercial emerging hardware.
Journal refRISC-V Summit Europe 2026 Bologna