AI 中文总结
研究低精度下矩阵和张量列的秩补偿问题,通过奇异值误差恒等式等方法,在SuiteSparse矩阵和多种张量测试中验证,实现存储减少且误差可控,不同精度有不同加速比,如矩阵中FP32和FP16分别有1.28倍和2.12倍加速,张量重建中分别有1.38倍和1.94倍加速。
AI 中文摘要
较低的数值精度可减少存储和内存流量,但会提高扰动下限。我们研究秩补偿,即将节省的内存重新投入到更大的近似秩中。对于矩阵,奇异值误差恒等式产生了一个可直接检验的充分条件,要求额外的奇异分量抵消因以较低精度存储秩增强近似而产生的扰动。在十个SuiteSparse矩阵上,所有100个截断主导配置(50个FP32和50个FP16)都被证明误差不增加且严格精度获胜,相对于FP64基线,平均误差率为0.963,存储率分别为58.8%和29.4%。FP16失败仅发生在接近扰动下限的尾秩压力测试中。在最大的常驻矩阵应用批次中,补偿后的FP32和FP16实现了几何平均A100加速比分别为1.28倍和2.12倍;两者都没有加速最小批次。对于张量列(TT)近似,我们基于测量的截断增益和舍入核心扰动给出了一个条件后验扩展。在三路和六路合成测试中,FP32和FP16在20次试验中的10次和14次试验中实现了精度 - 内存综合获胜。在公共高光谱张量和FROSTT顶级活跃子张量上,相应的计数分别为60次中的44次和60次中的54次;四个FP16萨利纳斯 - A尾应力情况失败。没有经过认证的TT情况超过FP64误差超出数值容差。六个公共张量的重建产生了几何平均补偿加速比分别为1.38倍(FP32)和1.94倍(FP16)。计时涵盖常驻下游内核,不包括因式分解、传输或端到端加速。
英文摘要
Lower numerical precision reduces storage and memory traffic but raises the perturbation floor. We study rank compensation: reinvesting saved memory in a larger approximation rank. For matrices, the singular-value error identity yields a directly testable sufficient condition requiring the additional singular component to offset the perturbation from storing the rank-augmented approximation in lower precision. On ten SuiteSparse matrices, all 100 truncation-dominated configurations (50 FP32 and 50 FP16) are certified non-increases and strict accuracy wins, with mean error ratio $0.963$ and storage ratios $58.8\%$ and $29.4\%$ relative to the FP64 baseline. FP16 failures occur only in tail-rank stress tests near the perturbation floor. At the largest resident matrix-application batch, compensated FP32 and FP16 achieve geometric-mean A100 speedups of $1.28\times$ and $2.12\times$; neither accelerates the smallest batch. For Tensor-Train (TT) approximation, we give a conditional a posteriori extension based on the measured truncation gain and rounded-core perturbation. Across three-way and six-way synthetic tests, FP32 and FP16 achieve combined accuracy-memory wins in 10 of 20 and 14 of 20 trials. On public hyperspectral tensors and FROSTT top-active subtensors, the corresponding counts are 44 of 60 and 54 of 60; four FP16 Salinas-A tail-stress cases fail. No certified TT case exceeds the FP64 error beyond numerical tolerance. Reconstruction of six public tensors yields geometric-mean compensated speedups of $1.38\times$ (FP32) and $1.94\times$ (FP16). Timings cover resident downstream kernels, not factorization, transfers, or end-to-end acceleration.
Comments24 pages, 5 figures, 19 tables