arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双四元数代数与双线性乘法:适用于CPU和GPU上相对论电子结构计算的通用、简洁且计算优势显著的框架

Biquaternion Algebra with Bilinear Multiplication: A General, Elegant, and Computationally Advantageous Framework for Relativistic Electronic Structure Calculations on CPUs and GPUs

Sylvia Kaviraj, Stanislav Komorovsky, Trond Saue, Michal Repisky

arXiv 2609.02081首次发表:更新:

发表机构

Comenius University in Bratislava; Slovak Academy of Sciences; Laboratoire de Chimie et Physique Quantiques, UMR 5626 CNRS — Université de Toulouse; UiT - The Arctic University of Norway(布拉迪斯拉发康斯坦丁神甫大学; 斯洛伐克科学院; 图卢兹大学量子化学与物理实验室,CNRS联合研究单位5626; 挪威阿尔塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文开发了双四元数矩阵库HMATLIB,提出双线性算法减少双四元数乘法的计算量,其在CPU和GPU上的性能优于复代数方法,为相对论电子结构计算提供了更优框架。

AI 中文摘要

四元数代数为相对论电子结构理论中时间反演对称的矩阵结构提供了自然表示,而互补的时间反演反对称结构则将该表示扩展到复四元数或双四元数。本文在统一框架内明确利用双四元数代数处理基本对象,包括算子、动力学平衡基函数和磁平衡基函数及其期望值,该框架涵盖实、复和实四元数子代数作为特例。为在计算上实现该框架,我们开发了HMATLIB,这是一个在ReSpect软件包内实现的双四元数矩阵库,支持纯CPU执行和CPU/GPU混合执行。一项关键进展是针对矩阵值双四元数乘法的双线性算法,与传统的分量式方法相比,该算法将实矩阵-矩阵乘法的数量从64次减少到24次,同时在相同的32位索引约束下,双四元数表示可有效将最大可访问矩阵维度加倍。在现代CPU和GPU架构上进行的数值基准测试表明,双四元数公式始终优于其同构复代数对应物。对于所报告的最大矩阵,CPU/GPU混合执行使双线性矩阵乘法相比纯CPU执行加速约5至8倍,矩阵对角化相比CPU oneAPI MKL加速约12倍,相比NVHPC OpenBLAS加速约61倍。这些结果确立了双四元数代数是现代相对论电子结构计算的通用且计算优势显著的框架。

英文摘要

Quaternion algebra provides a natural representation of time-reversal-symmetric matrix structures in relativistic electronic-structure theory, whereas complementary time-reversal-antisymmetric structures extend this representation to complex quaternions, or biquaternions. Here, we explicitly exploit biquaternion algebra for fundamental objects, including operators, kinetically and magnetically balanced basis functions, and their expectation values, within a unified framework encompassing real, complex, and real quaternion subalgebras as special cases. To realize this framework computationally, we have developed HMATLIB, a biquaternion matrix library implemented within the ReSpect package for pure CPU and hybrid CPU/GPU execution. A key development is a bilinear algorithm for matrix-valued biquaternion multiplication that reduces the number of real matrix--matrix multiplications from 64 to 24 compared with the conventional component-wise approach, while the biquaternion representation effectively doubles the maximum accessible matrix dimension under the same 32-bit indexing constraint. Numerical benchmarks on modern CPU and GPU architectures demonstrate that the biquaternion formulation consistently outperforms its isomorphic complex-algebra counterpart. For the largest matrices reported, hybrid CPU/GPU execution accelerates bilinear matrix multiplication by approximately 5-8 times over pure CPU execution and matrix diagonalization by approximately 12 and 61 times relative to CPU oneAPI MKL and NVHPC OpenBLAS, respectively. These results establish biquaternion algebra as a general and computationally advantageous framework for modern relativistic electronic-structure calculations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑