arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在CUDA上构建任意精度浮点运算环境及性能评估

Construction and Performance Evaluation of an Arbitrary-Precision Floating-Point Arithmetic Environment on CUDA

Tomonori Kouya

arXiv 2608.00085首次发表:更新:

AI 中文总结

本研究开发了适用于GPU的任意精度运算环境mpc_cuda,将MPFR/MPC库移植到CUDA内核,运算速度提升约40至105倍,在低至中等精度下优于现有GPU库CUMP和CAMPARY。

AI 中文摘要

任意精度浮点运算用于解决科学计算中的病态问题,MPFR和MPC已成为基于CPU计算的事实上的标准库。然而,目前没有公开可用的环境在GPU上提供MPFR/MPC功能。本研究利用生成式AI(Claude Code)和可重运行转换脚本,开发了适用于GPU的任意精度运算环境mpc_cuda,将基于mini-GMP的MPFR/MPC库忠实地移植到CUDA内核,同时保留每线程临时存储区域( bump arena)。该方法在GPU上的实型和复型初等函数及基础线性代数运算速度提升约40至105倍,且与主机结果位一致。此外,我们实现了编译时固定精度类型cu_freal/cu_fcomplex,通过在GB10和H100 GPU上进行的GMRES(m)方法及代数方程求解的基准测试表明,所提实现在低至中等精度下优于现有GPU库CUMP和CAMPARY。

英文摘要

Arbitrary-precision floating-point arithmetic is used to solve ill-conditioned problems in scientific computing, and MPFR and MPC have become the de facto standard libraries for CPU-based computations. However, no publicly available environment provides MPFR/MPC functionality on GPUs. In this study, we developed mpc\_cuda, an arbitrary-precision arithmetic environment for GPUs, by leveraging generative AI (Claude Code) and rerunnable conversion scripts to faithfully port the mini-GMP-based MPFR/MPC libraries to CUDA kernels while reserving a per-thread temporary storage region (a bump arena). This approach accelerates real and complex elementary functions and basic linear algebra on GPUs by approximately 40--105$\times$, while remaining bit-identical to the host. In addition, we implemented compile-time fixed-precision types cu\_freal/cu\_fcomplex, and demonstrate, through benchmark tests involving the GMRES($m$) method and algebraic equation solving on GB10 and H100 GPUs, that the proposed implementation outperforms the existing GPU libraries CUMP and CAMPARY at low-to-medium precision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑