发表机构
National Center for Supercomputing Applications, University of Illinois at Urbana-Champaign; Department of Mechanical Engineering, Stanford University(伊利诺伊大学厄巴纳-香槟分校国家超级计算应用中心; 斯坦福大学机械工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对照分析不同注意力机制的DeepONet变体,发现传感器级标记化结合交叉注意力可显著降低经典DeepONet的误差,查询依赖交叉注意力最可靠,分支自注意力适配大尺寸复杂函数输入。
AI 中文摘要
深度神经算子学习输入函数与完整偏微分方程(PDE)解场之间的映射,对新问题实例的正向评估速度比传统数值求解器快几个数量级。注意力机制最近被引入神经算子,但多数研究同时修改多个架构组件,难以确定实际提升精度的因素。本工作对5种具有不同注意力机制的深度算子网络(DeepONet)变体开展对照系统研究,这些变体在数据驱动和物理感知两种范式下训练,以分离交叉注意力、自注意力、标记化和注意力深度的影响。我们在三个基准问题上评估这些变体:源驱动的瞬态一维非线性扩散-反应方程、具有可变初始条件的瞬态一维粘性伯格斯方程,以及具有异构源场的二维泊松热传导问题。传感器级标记化结合交叉注意力,在所有基准训练组合中将经典DeepONet的平均相对L₂误差降低2.4至28.0倍,最优注意力配置则降低3.5至32.3倍。分支自注意力仅与点积融合配对时表现不一致,会降低一维问题的精度,但对更复杂的二维源场有帮助;在交叉注意力之上添加分支自注意力可改善所有6种情况,但提升幅度小于单独的交叉注意力融合。全局预混合无一致收益。增加交叉注意力深度可进一步提升精度,但收益递减且在物理感知训练下成本显著更高。总体而言,依赖查询的交叉注意力是最可靠的机制,而分支自注意力对大尺寸、空间复杂的函数输入最有用。
英文摘要
Deep neural operators learn mappings between input functions and complete PDE solution fields, enabling forward evaluations of new problem instances orders of magnitude faster than conventional numerical solvers. Attention mechanisms have recently been introduced into neural operators, but most studies change several architectural components at once, making it difficult to identify what actually improves accuracy. This work presents a controlled and systematic study of five deep operator network (DeepONet) variants with distinct attention mechanisms, trained under both data-driven and physics-informed regimes, to isolate the effects of cross-attention, self-attention, tokenization, and attention depth. We evaluate them on a source-driven transient one-dimensional nonlinear diffusion-reaction equation, a transient one-dimensional viscous Burgers equation with variable initial conditions, and a two-dimensional Poisson heat-conduction problem with heterogeneous source fields. Per-sensor tokenization with cross-attention reduces the mean relative L_2 error of the classical DeepONet in all benchmark-training combinations by factors of 2.4-28.0, while the best attention configurations reach 3.5-32.3. Branch self-attention paired only with dot-product fusion is inconsistent, degrading the one-dimensional problems while helping the more complex two-dimensional source field; added on top of cross-attention it improves all six cases, though by less than cross-attention fusion alone. Global pre-mixing provides no consistent benefit. Increasing cross-attention depth further improves accuracy, but with diminishing returns and a substantially higher cost under physics-informed training. Overall, query-dependent cross-attention is the most reliable mechanism, whereas branch self-attention is most useful for large, spatially complex functional inputs.
Comments25 pages, 13 figures, 5 tables