AI 中文总结
针对量子系统模拟对高维张量处理的需求,提出基于CUDA的雅可比型张量分解算法,在NVIDIA H100 GPU上实现9阶内可对角化张量的正确收敛,性能较MATLAB参考实现提升超两个数量级。
AI 中文摘要
量子系统的复杂性增长极快,即使是规模适中的模型也难以模拟,因此迫切需要能处理高维数据(即张量)的方法。本研究提出一种新型雅可比型张量分解算法,开发了基于CUDA的算法,支持单GPU上任意阶数的张量。在NVIDIA H100 GPU上测试该实现,结果表明该算法对最高9阶的可对角化张量能正确收敛,且运行时间随张量阶数增长呈可预测的缩放关系。最终,该通用算法的性能比原始MATLAB参考实现高出两个数量级以上。
英文摘要
Quantum systems grow in complexity so rapidly that even modest models become difficult to simulate, creating a strong need for methods that can handle high-dimensional data, also known as tensors. In this work, we investigate a novel Jacobi-type tensor algorithm for tensor decomposition and develop a CUDA-based algorithm that supports tensors of arbitrary order on a single GPU. We test the implementation on NVIDIA H100 GPUs and show that the algorithm converges correctly for diagonalizable tensors up to nine dimensions, with runtime scaling in a predictable way as tensor order grows. Finally, our general algorithm outperforms the original MATLAB reference by more than two orders of magnitude.