AI 中文总结
针对时谐麦克斯韦方程的极端规模求解难题,提出带PML和混合求积的完全无矩阵三网格预条件子,可高效求解百亿级未知量的均匀与异质系统,实现高可扩展性。
AI 中文摘要
三维时谐麦克斯韦模拟会产生庞大的复不定方程组,其网格粗化严格受相位精度限制。尽管无矩阵有限元内核能高效利用GPU吞吐量,但标准多级求解器最终会受限于精确粗网格分解的内存与通信成本。针对带完美匹配层(PML)和最优混合求积的旋度协调Nédélec离散格式,我们提出一种完全无矩阵、无分解的三网格预条件子。该方法采用外层FGMRES求解未偏移的细网格方程,同时通过固定工作量FGMRES计算中间网格校正项,该FGMRES由复偏移的2h-4h循环预条件化。此策略将复偏移限制在辅助预条件子中,保留物理麦克斯韦算子。局部傅里叶分析推导混合麦克斯韦分支与兼容边传递,确定稳健的偏移和雅可比阻尼参数。针对解析麦克斯韦格林张量的验证表明,该方法展现出极端可扩展性:采用单一求解器配置,在仅64个NVIDIA A100 GPU上,可在42.0-72.0秒内求解约108.9亿个复边未知量的均匀及高度异质系统。
英文摘要
Three-dimensional time-harmonic Maxwell simulations generate massive complex indefinite systems whose mesh coarsening is strictly limited by phase accuracy. Although matrix-free finite element kernels utilize GPU throughput efficiently, standard multilevel solvers are ultimately bottlenecked by the memory and communication costs of exact coarse-grid factorizations. We present a fully matrix-free, factorization-free three-grid preconditioner for curl-conforming N{é}delec discretizations with perfectly matched layers (PML) and optimally blended quadrature. The method employs an outer FGMRES to solve the unshifted fine-grid equation, while an intermediate-grid correction is computed by a fixed-work FGMRES preconditioned with a complex-shifted $2h$--$4h$ cycle. This strategically confines the complex shift to an auxiliary preconditioner, preserving the physical Maxwell operator. A local Fourier analysis derives the blended Maxwell branches and compatible edge transfers, identifying robust shift and Jacobi damping parameters. Validated against the analytical Maxwell Green tensor, our approach demonstrates extreme scalability: using a single solver configuration, both homogeneous and highly heterogeneous systems with approximately 10.89 billion complex edge unknowns are solved in 42.0--72.0 seconds on just 64 NVIDIA A100 GPUs.