arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30039physics.opticsphysics.comp-ph

TorchFDTD:用于光子逆向设计的GPU加速时域有限差分模拟,具有离散伴随和主机流式执行

Inverse design of large-scale freeform meta-optics by breaking the memory wall of full-wave simulation

Hyoseok Park, Yeonsang Park

首次发表
浏览论文内容

中文总结 AI 辅助

TorchFDTD通过CUDA图融合内核和离散伴随实现GPU加速FDTD,支持主机流式处理超大问题,在A100上比PyTorch FDTD快9-17倍,使数千万单元的光子逆向设计在单GPU上可行。

中文摘要 AI 辅助

基于梯度的光子设计需要对数百万个参数进行全波导数计算,但在单个GPU工作站上,时域伴随方法受到设备内存以及融合更新内核与自动微分不兼容的限制。本文介绍了TorchFDTD,一个开源时域有限差分(FDTD)软件包,解决了这两个限制。其Yee、吸收体和色散更新以融合CUDA内核形式执行,并捕获在CUDA图中,每个内核都配有一个由其更新推导出的转置内核,因此材料和几何导数通过离散伴随获得,PyTorch将其与基于支持观测构建的可微目标函数链接起来。伴随方法通过检查点重放或对于无损周期问题通过时间反转来恢复前向状态。对于超出设备内存的问题,流式模式每次推进一个因果切片,并将全局状态和检查点保存在主机内存中,从而降低设备内存分配,同时保留常驻离散化。我们针对解析解、Meep、FDTDX和严格耦合波求解器,以及自动微分和有限差分验证了该软件包。在A100上,融合路径完成八个测试场景的完整前向求解比其扩展的PyTorch FDTD软件包快9.0到17.1倍,其双精度求解比工作站CPU上的Meep快49到59倍。在RTX 3060上,$256^3$和$320^3$伴随的主机流式传输成本分别是常驻时间的3.1倍和2.7倍,并将峰值设备分配降低了56%和65%。一个5400万单元的柱阵列透镜耦合到角谱目标函数,产生的伴随导数与中心差分相差在0.78%以内。因此,具有数千万单元的器件的时域伴随在单个工作站GPU上变得可行。

英文摘要

Gradient-based design of large optical surfaces hits a memory wall: a stored-history time-domain adjoint needs $24N_t$ bytes per cell for $N_t$ steps, and every field update is bound by memory bandwidth. Here we present a validated framework that removes this wall and makes globally optimized freeform meta-optics practical. Its finite-difference time-domain adjoint reconstructs interior fields backwards from two recorded planes, so volume memory no longer grows with $N_t$, and fused kernels raise memory-bandwidth use on an H200 GPU from 6.4% to a remarkable 56% of peak. Overlapping tiles, coupled by angular-spectrum propagation and distributed over GPUs, keep tile memory independent of the aperture. Within hours on one GPU node, we optimize, against one global objective, every density variable of single-layer metasurfaces 100 to 200 $μ$m across, and design a nine-wavelength lens, a color hologram and a polarization-switched hologram. Confirmed by an independent solver within 1.6%, they substantially outperform meta-atom designs, reaching 1.64, 1.65 and 1.24 times their objectives.

发表机构

  • Chungnam National University(忠南国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑