arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向时变偏微分方程的基础模型蒸馏

Distillation of Foundation Models for Time-dependent PDEs

Daniel Musekamp, Boshra Ariguib, Andrei Manolache, Mathias Niepert

arXiv 2608.11937首次发表:更新:

发表机构

University of Stuttgart; University of Cologne; Bitdefender(斯图加特大学; 科隆大学; 比特 Defender(比特Defender,一家网络安全公司))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向时变PDE的知识蒸馏框架TREX,将预训练基础模型的预测能力迁移至紧凑高效的学生模型,在多PDE基准上实现精度相当或超越教师模型,且参数大幅减少、推理速度显著提升。

AI 中文摘要

面向时变偏微分方程(PDE)的基础模型在大量多样的物理系统集合上进行训练,能够有效泛化到新的下游任务;仅在目标领域的少量轨迹上微调后,它们就能在低数据 regime 中实现较高精度。然而,这些模型通常规模庞大且计算密集,限制了其作为数值求解器快速替代模型的实用性。我们提出 Teacher Rollout Extension(TREX),这是一种知识蒸馏框架,可将预训练基础模型的预测能力迁移至紧凑高效的学生模型。从微调后的教师模型出发,TREX 通过教师回滚生成长合成轨迹,可选地加入周期性噪声注入,以此扩充有限的下游数据;该过程从教师诱导的回滚分布中采样,无需明确知晓初始条件分布,同时让学生接触到长 horizon 状态以及自回归预测过程中遇到的状态周围的局部恢复行为。学生模型还可进一步融入教师未必强制的任务特定归纳偏置,比如等变性。我们在多个 PDE 基准上对 TREX 进行评估,所得学生模型的精度可与教师模型相当或超越,同时参数数量减少数个数量级,推理速度提升一个数量级以上。

英文摘要

Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few trajectories from a target domain, they can achieve strong accuracy in low-data regimes. However, these models are typically large and computationally intensive, limiting their usefulness as fast surrogates for numerical solvers. We propose Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student. Starting from a fine-tuned teacher, TREX augments limited downstream data by generating long synthetic trajectories through teacher rollouts, optionally with periodic noise injection. This procedure samples from the teacher-induced rollout distribution without requiring explicit knowledge of the initial-condition distribution, while exposing the student to long-horizon states and local recovery behavior around states encountered during autoregressive prediction. The student can further incorporate task-specific inductive biases, such as equivariance, that the teacher does not necessarily enforce. We evaluate TREX on multiple PDE benchmarks. The resulting students can match or surpass the teacher's accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑