发表机构
Faculty of Science and Engineering Otemon Gakuin University(大东文化大学理工学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对对算术中重归一化缺失导致迭代求解器不收敛的问题,提出无分支固定循环次数的重归一化方案RenormBF-pair,在K≤3时仅需一轮即可保证无重叠,集成到CG和BiCGStab中,在binary64下与VecSumK无差异,在binary32下显著减少违规,成本仅增加3.4%。
AI 中文摘要
对算术省略了多分量多精度算术每次操作结束时执行的重归一化阶段,从而减少了操作次数,并通过消除条件分支简化了SIMD向量化。然而,重归一化不能完全省略,因为迭代求解器可能因此无法收敛。我们提出了RenormBF-pair,一种无分支的固定循环次数重归一化器,以字数K和输入长度n为参数:一次VecSum扫描、一次尾部折叠以及r轮FastTwoSum链,代价为6(n-1)+(n-K)+3r(K-1)次浮点运算。在每种条件下超过200,000次试验中,产生无非重叠违规的最小轮数为:K≤3时r=1,K=4时r=2,而将相同代价用于VecSumK风格的TwoSum链在K=4时失败。我们通过显式反例界定了和关系成立时的输入条件,以及该方案通常不保证最低分量正确舍入的事实。将其集成到基于对算术的CG和BiCGStab中,在binary64和binary32下,我们进行了240次配对运行,包括从SuiteSparse筛选的真实矩阵。所提出的方案和VecSumK在两者都成功的地方返回了逐位相同的字,因此在binary64的CG中它们无法区分;差异仅出现在VecSumK失效的地方,即binary32和BiCGStab。在那里,22胜16负的记录并不显著(符号检验,p=0.42);明显的偏差是结构性的,非重叠违规从66行降至37行。整个CG求解器的成本增加了+3.4%。
英文摘要
Pair arithmetic omits the renormalization stage performed at the end of each operation of multi-component multiple-precision arithmetic, reducing the operation count and, by eliminating conditional branches, easing SIMD vectorization. Renormalization cannot be omitted entirely, however, because iterative solvers may then fail to converge. We propose RenormBF-pair, a branch-free fixed-trip-count renormalizer parameterized by the word count \(K\) and the input length \(n\): a VecSum sweep, a tail fold, and \(r\) rounds of a FastTwoSum chain, costing \(6(n-1)+(n-K)+3r(K-1)\) flops. Over \(200{,}000\) trials per condition, the smallest round count producing no non-overlap violation was \(r=1\) for \(K \leq 3\) and \(r=2\) for \(K=4\), whereas the same cost spent on a VecSum\(K\)-style TwoSum chain fails for \(K=4\). We delimit with explicit counterexamples both the input condition under which the sum relation holds and the fact that the scheme does not, in general, guarantee a correctly rounded lowest component. Integrating it into CG and BiCGStab over pair arithmetic in binary64 and binary32, we sweep \(240\) paired runs including real matrices screened from SuiteSparse. The proposed scheme and VecSum\(K\) returned bitwise identical words wherever both succeeded, so they are indistinguishable for CG in binary64; differences appear only where VecSum\(K\) breaks, namely binary32 and BiCGStab. There the \(22\)--\(16\) win--loss record is not significant (sign test, \(p=0.42\)); the clear bias is structural, non-overlap violations dropping from \(66\) rows to \(37\). The cost is \(+3.4\%\) for the whole CG solver.