arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解释自动编码器的学习动态:瞬态缩放和伊辛模型的新兴概念

Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

Max Weinmann, Miriam Klopotek

arXiv 2607.10285首次发表:更新:

发表机构

Stuttgart Center for Simulation Science, Cluster of Excellence EXC 2075, University of Stuttgart(斯图加特大学卓越模拟科学中心,卓越集群EXC 2075)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究伊辛模型微观自旋配置上训练的无监督自动编码器如何学习宏观变量,通过量化多尺度学习揭示两种动态模式,利用递归动态轨迹分析证明预测误差诱导流场,建立学习动态观点并提供解释基础。

AI 中文摘要

我们研究了在伊辛模型的微观自旋配置上训练的无监督自动编码器如何学习数据生成过程背后的宏观、与理论相关的变量。在不嵌入领域知识的情况下,我们模拟了一个典型的发现设置:量化跨多个空间(粗粒度)尺度的学习,并揭示由主要超参数(模型深度、宽度和学习率)控制的两种不同的动态模式——磁化主导模式和能量主导模式,其特征在于它们的表示质量存在权衡。第一种模式是一个过渡状态,表现出动态缩放和遵循有序到尺度的波动;第二种模式逐渐将分辨率转向与能量表示相关的较小尺度。以中等和快速速率训练的深度模型在达到这些模式之前就会停滞。通过对递归动态轨迹的新颖分析,我们证明预测误差会诱导流场,从而在所有表示空间中产生共同的轨迹拓扑。我们建立了一种学习的动态观点,其中内在属性揭示了训练期间表示强制变化所带来的影响。我们利用这样的直觉,即学习是一个由训练数据和优化器的波动驱动远离平衡的过程,从而提供一个基于物理世界和表示它的机器模型的解释基础。

英文摘要

We study how unsupervised autoencoders trained on microscopic spin configurations from the Ising model learn macroscopic, theory-relevant variables underlying the data-generating process. We quantify learning across multiple spatial (coarse-graining) scales and reveal two distinct dynamical regimes that appear sequentially, controlled by the main hyperparameters (model depth, width, and learning rate): one in which magnetization and another in which energy is learned across scales. The first exhibits error fluctuations ordered to scale and learns global averages only; The second gradually resolves smaller scales relevant for the energy representation. Deep models trained at moderate and fast rates become arrested before reaching these regimes. We connect reconstruction errors with the latent representations using a novel analysis of self-recursive trajectories. These intrinsic dynamics are induced by prediction errors, exposing how training drives representation changes for macroscopic concepts. We utilize the intuition that learning operates as a process driven far from equilibrium by fluctuations from the training data to provide an interpretive basis grounded in both the physical world and the machine models that represent it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑