AI 中文总结
该研究用Futhark实现GPU上中等规模大整数的块级四则运算,对比手写C++/CUDA与CGBN,发现数组自动布局对性能关键,高级代码经优化后性能具竞争力。
AI 中文摘要
我们采用高级函数式语言Futhark,实现了适用于GPU的块级加、减、乘、除运算,处理的操作数大小为2^15至2^19位的中等规模整数。通过与手写C++/CUDA版本及CGBN工具对比,明确了哪些函数式结构可顺利编译、内存布局与序列化策略的有效场景,以及所需的编译器支持。结果表明,高级代码能紧凑表达算法,经编译器优化后可达到有竞争力的性能,其中GPU寄存器内存中数组的自动布局对性能至关重要。
英文摘要
We report on GPU implementations of block-level addition, subtraction, multiplication and division for midsize integers, with operands of $2^{15}$ to $2^{19}$ bits using the high-level functional language Futhark. Comparing with hand-written C++/CUDA versions and CGBN, we identify which functional constructs compile well, where memory placement and sequentialization are effective, and what compiler support is needed. The results show that high-level code can express the algorithms compactly while approaching competitive performance after certain compiler improvements. In particular, we find that automated placement of arrays in GPU register memory is critical for performance.