arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38814cs.LGnlin.CD

无见混沌而学混沌:自回归Transformer中全局动力学的外推

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明,仅从受限参数区域的局部轨迹训练的小型自回归Transformer,能在未见参数下高保真外推混沌系统的全局动力学,如逻辑斯蒂映射重现倍周期级联至周期128,标度比4.6687接近费根鲍姆常数,并揭示注意力机制在其中的作用。

中文摘要 AI 辅助

自回归模型被训练用于逐步预测系统的行为,递归生成使得学习到的动力学能够在长时间范围内展开。从局部观测中学到的这种动力学,在多大程度上能够恢复训练期间仅部分观测到的底层系统的更广泛组织?在此,我们研究在若干非线性动力系统的受限参数区域采样的轨迹上从头训练的小型自回归Transformer,包括逻辑斯蒂映射、正弦映射、洛伦兹系统以及广义霍普夫系统,其中控制参数和状态轨迹表示为连续令牌序列。在训练分布之外的参数下进行闭环评估时,模型能够以显著的视觉和数值保真度恢复自相似的倍周期级联、混沌动力学和吸引子结构。对于逻辑斯蒂映射,一个Transformer重现了直至周期128的连续倍周期,产生有限阶标度比4.6687,与费根鲍姆常数在$5\ imes10^{-4}$以内吻合。我们进一步研究了这些结构在训练过程中如何涌现,并通过因果干预揭示了控制参数信息如何通过注意力被处理到状态预测中,从而塑造所得的闭环动力学。这些结果表明,对系统局部行为的一个惊人狭窄窗口可能足以让自回归Transformer泛化到其未见过的全局动力学组织。

英文摘要

Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed during training? Here we study small autoregressive transformers trained from scratch on trajectories sampled from restricted parameter regimes of several non-linear dynamical systems, including logistic and sine maps, the Lorenz system, and the generalized Hopf system, with control parameters and state trajectories represented as sequences of continuous tokens. Under closed-loop evaluation at parameters far outside the training distribution, the models can recover self-similar period-doubling cascades, chaotic dynamics, and attractor structures with remarkable visual and numerical fidelity. For the logistic map, a transformer reproduces successive period doublings up to period 128, yielding a finite-order scaling ratio of 4.6687, matching the Feigenbaum constant to within $5\times10^{-4}$. We further investigate how these structures emerge over the course of training, and reveal with causal interventions how control-parameter information is processed through attention into state prediction and shapes the resulting closed-loop dynamics. These results suggest that a surprisingly narrow window into a system's local behavior may suffice for autoregressive transformers to generalize to its unseen global dynamical organization.

发表机构

  • Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)
  • Munich Center for Machine Learning(慕尼黑机器学习中心)
  • Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑