arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPEAR:一种用于大规模PDE预训练的知识引导专家聚合的谱解耦MoE神经算子

SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining

Dengdi Sun, Xiaoya Zhou, Xiao Wang, Wanli Lyu, Jin Tang, Bin Luo

arXiv 2610.03265首次发表:更新:

发表机构

Anhui University(安徽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对PDE预训练中异质动力学和专家冗余问题,提出谱解耦MoE神经算子SPEAR,通过知识引导聚合减少50%专家并保持精度,在12个数据集上验证了优越性能。

AI 中文摘要

大规模预训练提升了神经算子在不同PDE上的泛化能力。然而,现有的PDE基础模型在处理异质动力学时仍面临挑战,其中共享表示可能导致知识干扰,而混合专家(MoE)架构则遭受专家冗余增加的问题。我们提出SPEAR,一种用于大规模PDE预训练的谱解耦MoE神经算子,具有知识引导的专家聚合功能。SPEAR将潜在特征解耦为低频和高频分量,从而实现对可迁移动力学的共享建模和对PDE特定模式的专门学习。为解决专家冗余问题,我们设计了一种知识引导的专家聚合策略,该策略根据数据集特定的学习知识和路由偏好来衡量专家相似性,从而能够识别和合并相似的专家。在十二个PDE数据集和多个下游基准上的实验表明,SPEAR在预训练、微调和迁移学习中均表现出优越的性能。此外,我们的聚合策略将专家数量减少了50%,同时保持或提高了预测精度,为PDE基础模型在模型效率和泛化之间取得了平衡。

英文摘要

Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50\% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑