Latent-MoE:面向多区域物理偏微分方程的域感知混合专家模型
Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对物理随区域变化的PDE,提出Latent-MoE,通过域感知MoE块和共享骨干实现局部化学习,在多阶段时变物理基准上性能提升超一个数量级。
中文摘要 AI 辅助
物理信息神经网络(PINNs)在控制物理随区域变化的偏微分方程(PDEs)上表现不佳。我们将此归因于标准坐标网络的一个结构特性:其神经正切核(NTK)是平移变化的,使得坐标幅值较大的训练点不成比例地影响其他位置的预测,从而在训练过程中产生长程耦合和梯度冲突。我们在分析和实证上表明,采用中心化、紧支撑路由器的混合专家(MoE)架构能够产生均匀带状NTK,其核回归权重随距离指数衰减,从而将学习局部化。基于此,我们提出Latent-MoE,在共享骨干网络中交错插入域感知的MoE块。与FB-PINNs或X-PINNs不同,后者严格划分域和参数,使得不同子域上的参数独立更新,Latent-MoE旨在保留域感知路由的局部化优势,同时允许能力通过共享骨干网络在区域间流动。在标准均匀物理基准上,Latent-MoE与既有基线相当;在具有多阶段时变物理的基准上,全局模型和刚性域分解均陷入伪解,而Latent-MoE将其性能提升超过一个数量级,且训练过程中的梯度冲突显著减少。
英文摘要
Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. We trace this to a structural property of standard coordinate networks: their neural tangent kernel (NTK) is translation-variant and lets training points of large coordinate magnitude disproportionately influence predictions elsewhere, producing long-range coupling and gradient conflict during training. We show analytically and empirically that mixture-of-experts (MoE) architectures with centered, compact-support routers yield a uniformly banded NTK whose kernel-regression weights decay exponentially with distance, localizing the learning. Building on this, we propose \emph{Latent-MoE}, which interleaves domain-aware MoE blocks within a shared backbone. Unlike FB-PINNs or X-PINNs, which rigidly partition both the domain and the parameters so that the parameters on different subdomains are updated independently, Latent-MoE is designed to preserve the localization benefit of domain-aware routing while allowing capacity to flow across regions through the shared backbone. On standard homogeneous-physics benchmarks Latent-MoE is competitive with established baselines; on benchmarks with multi-stage time-variable physics, where global models and rigid domain decompositions both fall into spurious solutions, it improves over them by more than an order of magnitude, with markedly reduced gradient conflict during training.