AI 中文总结
PhyMo提出物理场模态,通过PDE关联算子组织异构测量,经三阶段学习在五个数据集上超越最强基线,实现AI4Physics多模态表示学习的最优性能。
AI 中文摘要
多模态学习正成为AI for Physics(AI4Physics)中的一种强大范式,其中预测物理系统需要对异构观测、测量和领域知识进行联合解释。然而,现有方法通常将物理量和控制方程表示为通用数值或文本标记,忽略了决定其时空相互作用的物理约束。为解决这一局限性,我们引入了物理场模态,并提出了PhyMo,一个基于物理的多模态框架,通过PDE关联算子组织异构测量。PhyMo遵循三阶段学习流程:首先通过PDE残差监督下的场重建对物理场编码器进行预训练,随后将其表示与共享潜在空间中的视觉嵌入对齐,最后将融合的多模态表示交由相应的下游预测头处理。在涵盖不同物理环境的五个数据集上的实验表明,PhyMo在每个数据集上相比最强基线均取得了最先进的性能,证明了PhyMo在AI4Physics中多模态表示学习上的优越性。
英文摘要
Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.
CommentsUnder review