发表机构
University of California, Santa Cruz; University of Chicago; International Computer Science Institute; Lawrence Berkeley National Laboratory; University of California, Berkeley(加州大学圣克鲁兹分校; 芝加哥大学; 国际计算机科学研究所; 劳伦斯伯克利国家实验室; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出特征分析框架揭示神经自回归混沌模拟器误差增长的动力学起源,设计促进稳定性的损失函数,提升模型精度与鲁棒性,为混沌模拟器研发提供理论基础。
AI 中文摘要
神经自回归模型已迅速成为高维混沌系统的强大模拟器,但其长期不稳定性与误差增长机制仍未被充分理解,导致研究人员采用临时解决方案。本文中,我们开发了一种特征分析框架,揭示了这种误差增长的动力学起源。通过分析学习到的单步更新映射相对于状态的雅可比矩阵,我们证明了推理时的误差增长(即模型稳定性)如何由其谱半径决定。直接步架构(即从前一状态预测下一状态的模型)通常存在幅度超过1的不稳定特征值,这解释了这些广泛使用的模型的快速发散。相比之下,积分约束模型(即估计时间导数并使用高阶积分器进行积分的模型)会将其特征谱坍缩到单位圆上,产生中性稳定性和通用线性误差缩放定律。该雅可比矩阵的最大特征值提供了一种与架构无关的先验诊断方法,可用于判断短期技能、长期稳定性和谱偏差,而无需进行代价高昂的滚动推演。利用该理论,我们引入了一种促进稳定性的损失函数,该函数明确正则化由雅可比矩阵驱动的误差放大,从而提高了预测准确性和动力学鲁棒性。我们在Kuramoto-Sivashinsky系统上对29个模型(涵盖两种架构、多种显式和隐式积分器以及多种损失函数)进行了验证,结果为混沌多尺度动力学神经模拟器的设计与评估奠定了理论基础。更广泛地说,我们的框架朝着科学机器学习目前所缺乏的、数值分析为微分方程离散化提供的那种先验稳定性分析迈出了一步。
英文摘要
Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instability and error growth remain poorly understood, leading to ad-hoc solutions. Here, we develop an eigenanalysis framework that reveals the dynamical origin of this error growth. By analyzing the Jacobian of the learned one-step update map with respect to the state, we show how inference-time error growth, and thus model stability, is governed by its spectral radius. Direct-step architectures (models that predict the next state from the previous one) generically admit unstable eigenvalues with magnitudes exceeding one, explaining the rapid divergence of these widely used models. In contrast, integration-constrained models (where the time derivative is estimated and integrated with a higher-order integrator) collapse their eigenspectrum onto the unit circle, yielding neutral stability and a universal linear error-scaling law. The largest eigenvalue of this Jacobian provides an architecture-agnostic, a priori diagnostic of short-term skill, long-term stability, and spectral bias, without requiring an expensive rollout. Leveraging this theory, we introduce a stability-promoting loss that explicitly regularizes Jacobian-driven error amplification, improving both forecast accuracy and dynamical robustness. Demonstrated across $29$ models spanning two architectures, several explicit and implicit integrators, and multiple loss functions on the Kuramoto-Sivashinsky system, our results establish a theoretical foundation for the design and evaluation of neural emulators of chaotic multi-scale dynamics. More broadly, our framework is a step toward the kind of a priori stability analysis that numerical analysis provides for discretizations of differential equations and that scientific machine learning currently lacks.