浮点归约树的二阶矩理论
A Second-Moment Theory for Floating-Point Reduction Trees
浏览论文内容
中文总结 AI 辅助
研究浮点归约树求和误差与部分和顺序的关系,推导精确MSE递推公式,用统计量表征最优拓扑和调度,通过扩展核到矩阵乘法进行测试,模型能恢复拓扑排序,但低精度和的问题限制了其适用性。
中文摘要 AI 辅助
求和误差取决于部分和的顺序,而标准最坏情况界限忽略了这一点。为了捕捉这种依赖性,我们在条件无偏舍入下推导了二叉归约树\(T\)的精确均方误差(MSE)递推公式。在单位舍入\(u\)下,常数\(\nu\)模型将舍入前值\(x\)处的局部方差设置为\(\nu u^2 x^2\)。其对于输入向量\(p\)的主要树相关成本是\(p^T K_T p\),其中公共祖先核\(K_T\)计算每对叶子共享的内部祖先数量。对于均值为\(\mu\)、方差为\(\tau^2\)的独立同分布输入,此预期成本为\(\tau^2 \Lambda_1(T) + \mu^2 \Lambda_2(T)\),其中\(\Lambda_1\)是总叶深度,\(\Lambda_2\)是内部子树大小平方的总和;\(\Lambda_1\)控制中心化输入,而\(\Lambda_2\)捕捉非零均值。我们使用这些统计量来表征最优树拓扑和调度。平衡树和顺序树达到中心化极值。对于\(k\)个输入,最优两阶段顺序分块产生的均方根(RMS)误差缩放为\(k^{3/4}\)。对于固定阶段层次结构,几何调度对于中心化输入是最优的,而最优非中心化阶段指数依次减半。对于具有不等方差的独立中心化输入,霍夫曼编码在自由叶分配上最小化方差加权深度。我们通过操作数Gram矩阵将核扩展到矩阵乘法。然后我们使用精确残差测试在向最接近值舍入下的近似。在二进制64、二进制32以及软件模拟的二进制16和bfloat16中,该模型恢复了树拓扑之间的排序;\(K_T\)跟踪AR(1)部分和成本。对于通用矩阵乘法(GEMM),在测试网格上独立校准的预测与测量值的差异最多为3%。从数组库中提取的归约树预测了测量的RMS缩放。然而,正低精度和中的停滞和偏差限制了该模型的适用性。
英文摘要
Summation error depends on partial-sum order, which standard worst-case bounds omit. To capture this dependence, we derive an exact mean-square error (MSE) recurrence for a binary reduction tree T under conditionally unbiased rounding. With unit roundoff u, the constant-nu model sets the local variance at pre-rounding value x to nu u^2 x^2. Its leading tree-dependent cost for the input vector p is p^T K_T p, where the common-ancestor kernel K_T counts the internal ancestors shared by each pair of leaves. For i.i.d. inputs of mean mu and variance tau^2, this expected cost is tau^2 Lambda_1(T) + mu^2 Lambda_2(T), where Lambda_1 is total leaf depth and Lambda_2 sums squared internal-subtree sizes; Lambda_1 governs centered inputs, while Lambda_2 captures nonzero means. We use these statistics to characterize optimal tree topologies and schedules. Balanced and sequential trees attain the centered extrema. For k inputs, optimal two-stage sequential blocking yields root-mean-square (RMS) error scaling as k^{3/4}. For fixed-stage hierarchies, geometric schedules are optimal for centered inputs, whereas the optimal noncentered stage exponents halve successively. For independent centered inputs with unequal variances, Huffman coding minimizes variance-weighted depth over free leaf assignments. We extend the kernel to matrix multiplication through operand Gram matrices. We then test the approximation under round-to-nearest using exact residuals. Across binary64, binary32, and software-emulated binary16 and bfloat16, the model recovers the ordering among tree topologies; K_T tracks AR(1) partial-sum costs. For GEMM, independently calibrated predictions differ from measurements by at most 3% on the tested grid. A reduction tree extracted from an array library predicts the measured RMS scaling. However, stagnation and bias in positive low-precision sums limit the model's applicability.