发表机构
Google LLC(谷歌有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种仅用浮点运算、无分支的 Payne-Hanek 变体,通过截断误差预算和紧凑查找表,实现大双精度输入的高精度快速区间缩减,并已在 LLVM libc 中实现。
AI 中文摘要
区间缩减在三角函数求值的精度和性能中起着关键作用,并且常常是处理大浮点数输入时的主要瓶颈。虽然诸如 Cody--Waite 之类的快速算法在窄区间内能高效工作,但 Payne--Hanek 算法仍然是针对大浮点数输入进行精确缩减的标准技术。然而,现有的 Payne--Hanek 实现由于大量分支、转换开销以及使用多字整数运算而遭受高延迟,这阻碍了 SIMD 向量化。在本文中,我们分析并提出了一种仅使用浮点运算的 Payne--Hanek 算法的无分支变体。我们的方法直接处理大的双精度输入($|x| \ge 2^{16}$),并且非常适合具有 FMA 指令的硬件。我们根据截断误差预算来表述精度约束,构建一个由输入指数索引的紧凑查找表,并证明对于每个输入,缩减后的自变量精度在一个 ulp 以内。同一例程既可以作为单级实现的完整区间缩减,也可以作为正确舍入实现的快速路径,并且它比现有实现提高了延迟和吞吐量。该算法目前已在 LLVM libc 项目中实现。
英文摘要
Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, many existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs ($|x| \ge 2^{16}$) and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the scaled reduced argument has absolute error below $2^{-110}$ and relative error below $2^{-60}$ for every input, including worst cases. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, achieving higher throughput than existing implementations and lower latency than those returning a double-double reduced argument. The algorithm is currently implemented in the LLVM libc project.
CommentsUpdated to include CORE-MATH newer Payne-Hanek range reduction implementation, and add CRLIBM to performance analysis. Updated the performance table after fixing the setup