arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种加速张量分解的高性能计算(HPC)方法

An HPC Approach to Accelerate Tensor Decompositions

Markus Hellgren, Erna Begovic Kovac, Hans O. Karlsson, Roman Iakymchuk

arXiv 2608.24307首次发表:更新:

AI 中文总结

针对量子系统模拟对高维张量处理的需求,提出基于CUDA的雅可比型张量分解算法,在NVIDIA H100 GPU上实现9阶内可对角化张量的正确收敛,性能较MATLAB参考实现提升超两个数量级。

AI 中文摘要

量子系统的复杂性增长极快,即使是规模适中的模型也难以模拟,因此迫切需要能处理高维数据(即张量)的方法。本研究提出一种新型雅可比型张量分解算法,开发了基于CUDA的算法,支持单GPU上任意阶数的张量。在NVIDIA H100 GPU上测试该实现,结果表明该算法对最高9阶的可对角化张量能正确收敛,且运行时间随张量阶数增长呈可预测的缩放关系。最终,该通用算法的性能比原始MATLAB参考实现高出两个数量级以上。

英文摘要

Quantum systems grow in complexity so rapidly that even modest models become difficult to simulate, creating a strong need for methods that can handle high-dimensional data, also known as tensors. In this work, we investigate a novel Jacobi-type tensor algorithm for tensor decomposition and develop a CUDA-based algorithm that supports tensors of arbitrary order on a single GPU. We test the implementation on NVIDIA H100 GPUs and show that the algorithm converges correctly for diagonalizable tensors up to nine dimensions, with runtime scaling in a predictable way as tensor order grows. Finally, our general algorithm outperforms the original MATLAB reference by more than two orders of magnitude.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑