用于多分量类型多精度浮点运算的无分支融合乘加算法的性能评估
Performance evaluation of branch-free fused multiply-add algorithms for multi-component-type multiple-precision floating-point arithmetic
浏览论文内容
中文总结 AI 辅助
研究多分量多精度浮点运算的无分支融合乘加算法性能,基于Zhang和Aiken的工作,实现TW和QW变体并证明加速效果,提出新操作,经测试可进一步提升性能。
中文摘要 AI 辅助
在由现有浮点运算通过无误差变换(EFTs)构建的多分量多精度运算中,使用消除if分支的无分支算法可实现高性能。Zhang和Aiken提出了双字(DW)、三字(TW)和四字(QW)运算中加法和乘法的无分支算法。本文实现了TW和QW变体并证明其加速效果,还提出新的DW、TW和QW运算的无分支融合乘加操作,经基准测试表明能提升性能。
英文摘要
Multicomponent multiple-precision arithmetic is constructed from existing floating-point operations using error-free transformations (EFTs). Its performance can be improved by employing branch-free algorithms that eliminate conditional branches. Zhang and Aiken proposed branch-free addition and multiplication algorithms for double-word (DW), triple-word (TW), and quad-word (QW) arithmetic. We implemented their TW and QW algorithms, for which substantial performance improvements over conventional algorithms were expected, and demonstrated their effectiveness. In this paper, we propose branch-free fused multiply-add (FMA) algorithms for DW, TW, and QW arithmetic. The proposed algorithms integrate multiplication and addition into a single computational network and require fewer arithmetic operations than separately performing branch-free multiplication and addition. The anchor-relative error bounds, the preconditions of all FastTwoSum operations, and the non-overlapping properties of the outputs are mechanically verified using FPANVerifier, and input-relative error bounds are subsequently derived analytically. Benchmark results on CPUs and GPUs show that the proposed algorithms provide performance improvements in many compute-intensive cases, including division, square root, and basic linear algebra kernels, while maintaining accuracy comparable to that of the existing branch-free algorithms.