arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer的平均场理论:耦合数据-参数动力学的适定性与训练的全局收敛性

A Mean-Field Theory of Transformers: Well-Posedness of the Coupled Data--Parameter Dynamics and Global Convergence of Training

Michael Herty, Hailiang Liu

arXiv 2608.25055首次发表:更新:

AI 中文总结

该研究构建Transformer的平均场理论,分析耦合数据-参数动力学的适定性,证明浅层注意力模型指数收敛、深层Transformer局部线性收敛,为Transformer训练提供理论基础。

AI 中文摘要

我们为Transformer网络构建了一套严谨的平均场理论,该理论捕捉了架构固有的两个大规模极限:输入序列中的token数量$N\to\infty$,以及每一层中的注意力头数量$H\to\infty$。它还考虑了无限多层,进而得到时间连续的公式。所得框架耦合了两个相互作用的平均场对象:token分布$\mu_t\in\mathcal P(\mathbb R^d)$,其随网络深度$t\in[0,T]$通过McKean-Vlasov输运方程演化;注意力参数分布$\rho_s\in\mathcal P(\Theta)$,其随训练时间$s\ge0$通过经验风险的Wasserstein梯度流演化,可选择加入熵正则化或Tikhonov正则化。我们为耦合Transformer与训练动力学的系统建立了全面的分析基础,尤其确定了描述耦合平均场与训练动力学的非线性Fokker-Planck系统的全局适定性。除适定性外,我们还将平均场公式与优化相联系:对于浅层单注意力层模型,在对数Sobolev条件下,我们证明其指数收敛到熵正则化的全局最优解;对于真正深层的组合式Transformer,在神经正切核非退化条件下,我们建立了局部线性收敛性。

英文摘要

We develop a rigorous mean-field theory for transformer networks that captures two large-scale limits inherent in the architecture: the number of tokens $N\to\infty$ in the input sequence and the number of attention heads $H\to\infty$ in each layer. It further considers an infinite number of layers leading to a time-continuous formulation. The resulting framework couples two interacting mean-field objects: a token distribution $μ_t\in\mathcal P(\mathbb R^d)$, which evolves through network depth $t\in[0,T]$ according to a McKean--Vlasov transport equation, and the attention-parameter distribution $ρ_s\in\mathcal P(Θ)$, that evolves through training time $s\ge0$ according to a Wasserstein gradient flow of the empirical risk, with optional entropic or Tikhonov regularization. We establish a comprehensive analytical foundation for the system coupling the transformer and the training dynamics; in particular, we establish global well-posedness of the resulting nonlinear Fokker--Planck system describing the coupled mean-field and training dynamics. Beyond well-posedness, we connect the mean-field formulation to optimization. For shallow, single-layer attention models, we prove exponential convergence to the entropy-regularized global optimum under a log-Sobolev condition. For genuinely deep, compositional transformers, we establish local linear convergence under a Neural Tangent Kernel non-degeneracy condition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑