AI 中文总结
该研究针对Transformer自注意力栈的秩双指数衰减问题,提出用逐token黎曼度量替代欧几里得度量的理论框架,设计纤维丛Transformer架构,推导相关理论预测并明确核心开放问题。
AI 中文摘要
所有基于Transformer的大语言模型均通过欧几里得内积计算注意力,Dong等人(2021)证明,在纯自注意力栈中,该架构选择会导致表示秩随深度呈双指数衰减。我们开发了一个理论框架,从数学层面解决这一结构局限性,方法是将平坦的欧几里得度量替换为学习得到的逐token黎曼度量。我们的贡献有三:(1)证明具有异构逐token度量的黎曼注意力分数是非Gram矩阵——无法分解为QK^T形式,且分解维度为O(d),我们明确这是结构观察,而非秩保持的证明;(2)证明低秩度量因子使所有几何运算可处理:每个token的测地距离计算复杂度为O(d*r),通过Woodbury恒等式进行的度量求逆复杂度为O(d*r²),均远低于一般矩阵的O(d³)成本,使黎曼注意力可在十亿参数规模下运行,且开销可忽略;(3)提出纤维丛Transformer(Fiber Bundle Transformer),这是一个完整的架构规范,其中每个token位置携带自身的黎曼度量,注意力为测地距离计算,前馈更新使用度量预条件步骤,连接携带显式曲率和挠率代理。我们推导了正确实现的几何架构的形式预测,并确定核心开放问题:证明或证伪异构黎曼度量是否能防止行随机注意力矩阵否则会导致的秩崩溃。本文呈现理论分析与架构设计;实验验证是未来工作的主题。
英文摘要
All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank to decay doubly exponentially with depth in pure self-attention stacks. We develop a theoretical framework that targets this structural limitation at the mathematical level by replacing the flat Euclidean metric with learned per-token Riemannian metrics. Our contributions are threefold. (1) We prove that Riemannian attention scores with heterogeneous per-token metrics are non-Gram---they cannot be factorized as QK^T with factorization dimension O(d). We are explicit that this is a structural observation, not a proof of rank preservation. (2) We establish that low-rank metric factors render all geometric operations tractable: geodesic distance in O(d*r) per token and metric inversion in O(d*r^2) via the Woodbury identity---both far below the O(d^3) cost of a general matrix---making Riemannian attention feasible at billion-parameter scale with negligible overhead. (3) We present the Fiber Bundle Transformer, a complete architecture specification in which each token position carries its own Riemannian metric, attention is geodesic distance computation, feed-forward updates use metric-preconditioned steps, and the connection carries explicit curvature and torsion proxies. We derive formal predictions about correctly implemented geometric architectures and identify the central open problem: proving or disproving that heterogeneous Riemannian metrics prevent the rank collapse that row-stochastic attention matrices otherwise cause. This paper presents theoretical analysis and architectural design; empirical validation is the subject of future work.
Comments23 pages, theoretical paper