发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无重建学习潜在动力学中的尺度坍缩问题,提出交替训练的无重建框架,理论给出收敛条件,实验在三个任务上显著提升宏观预测。
AI 中文摘要
对复杂系统宏观性质的时间演化进行建模是一项重要的科学任务。为了在不进行完整微观模拟的情况下预测这种演化,一种常见的方法是将微观状态编码为紧凑的潜在状态,学习其演化,并从潜在轨迹中读出宏观预测。这些潜在状态通常通过微观状态重建来学习。然而,在潜在容量有限的情况下,重建可能偏向于高方差的微观细节,而非宏观预测所需的信息。然而,在不进行重建的情况下联合学习潜在状态及其转移,往往无法获得支持准确宏观预测的潜在动力学。我们表明,这种失败可能源于潜在尺度坍缩:缩小潜在状态尺度会降低训练损失,而宏观演化误差仍然很大。在此,我们提出一个无重建框架,用于学习具有其动力学的潜在状态,以实现指定的宏观预测。训练在更新潜在表示(固定转移和下一状态潜在目标)与更新转移(固定潜在表示)之间交替进行。在推理时,训练好的模型从初始微观状态递归地预测宏观状态。我们的理论分析刻画了联合训练下重建错位和尺度坍缩的特征,并给出了我们方法局部收敛到正确潜在动力学的充分条件。在晶格上的流行病传播、两种粒子物种的混合以及聚合物拉伸实验表明,所提出的方法在宏观预测方面显著优于基线方法。
英文摘要
Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolution, and reads out macroscopic predictions from the latent trajectory. These latent states are often learned through microstate reconstruction. However, with limited latent capacity, reconstruction can favor high-variance microscopic details over information needed for macroscopic prediction. Yet jointly learning latent states and their transition without reconstruction often fails to obtain latent dynamics that support accurate macroscopic prediction. We show that this failure can arise from latent scale collapse: shrinking the latent state scale reduces training loss while macroscopic evolution error remains large. Here, we propose a reconstruction-free framework to learn latent states with their dynamics for prescribed macroscopic prediction. Training alternates between updating the latent representation with the transition and next-state latent targets fixed, and updating the transition with the latent representation fixed. At inference, the trained model predicts macroscopic states recursively from an initial microstate. Our theoretical analysis characterizes reconstruction misalignment and scale collapse under joint training, and gives a sufficient condition for local convergence to correct latent dynamics for our method. Experiments on epidemic spreading on a lattice, mixing of two particle species, and polymer stretching demonstrate that the proposed method achieves substantially better macroscopic prediction over baselines.
CommentsWithdrawn by the authors due to concerns regarding the timing of public dissemination of results based on a dataset used in the paper.