AI 中文总结
研究Transformer动力学,构建几何框架,将令牌动力学建模为粒子系统。分离出建模语言所需的两个特征,证明Transformer模型能同时实现。给出一族通用的有限参数化相互作用定律,表明实现代价由有向图的组合不变量决定。
AI 中文摘要
我们开发了一个几何框架,其中Transformer的令牌动力学由黎曼流形$\mathcal M$上的相互作用粒子系统建模,注意力机制由一个与时间无关的两体相互作用定律编码,即$\mathcal M\times\mathcal M$上拉回丛$\pi_2^{*}(T\mathcal M)$的一个截面。在此框架内,我们分离出一族相互作用定律为建模语言必须具备的两个特征:它必须实现一般的非局部和非互易力,且必须有效地对高维流形上的向量场进行参数化。我们表明在Transformer模型中这两个特征能同时实现。我们的主要定理产生了一族有限参数化的相互作用定律,与流形及其维度无关,且具有通用性:它能实现任意规定的注意力有向图。此外,我们表明实现给定注意力有向图的代价不由$\dim\mathcal M$决定,而是由该有向图的两个组合不变量决定,即其二分图覆盖数(我们将其与中心扩展中的最少中心数等同)及其中心色指数。
英文摘要
We develop a geometric framework in which the token dynamics of a transformer are modeled by a system of interacting particles on a Riemannian manifold $\mathcal M$, the attention mechanism being encoded by a time-independent two-body interaction law, that is, a section of the pullback bundle $π_2^{*}(T\mathcal M)$ over $\mathcal M\times\mathcal M$. Within this framework we isolate two features that a family of interaction laws must possess in order to model language: it must realize generic nonlocal and nonreciprocal forces, and it must parametrize vector fields on a high-dimensional manifold efficiently. We show that both features are achieved simultaneously in a transformer model. Our main theorem produces a finitely parametrized family of interaction laws that is \emph{universal}: it realizes an arbitrary prescribed attention digraph. Moreover, we show that the cost of realizing a given attention digraph is governed by two combinatorial invariants of the digraph, namely its biclique cover number, which we identify with the least number of hubs in a hub extension, and its hub-chromatic index.