arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从自注意力到连接拉普拉斯算子:Transformer的统一算子视角

From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers

Binbin Lin, Wei Chen, Yalun Li, Wenxiao Wang, Jieping Ye, Xiaofei He

arXiv 2607.10677首次发表:更新:

发表机构

Zhejiang University; Alibaba Cloud(浙江大学; 阿里云)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从算子视角理解自注意力,将令牌序列视为向量场,证明单头和多头注意力的性质及相关条件,通过实验发现不同规模和结构的训练Transformer呈现与理论一致的几何结构,提供连接自注意力与几何算子的形式及分析工具。

AI 中文摘要

自注意力是现代序列模型中普遍存在的原语,但其算子级几何仅部分被理解。我们将令牌序列视为令牌位置图上的向量场,并将注意力识别为连接游走:消息由非负游走矩阵聚合,同时通过学习的线性映射沿每条边传输。在此框架内,我们证明单头注意力(SHA)恰好是具有恒定传输的连接传播步骤,多头注意力(MHA)恰好是单个依赖边的连接游走,其有效传输是各头传输的注意力门控混合。我们进一步阐明了相应生成器简化为随机游走连接拉普拉斯算子的条件,突出了随机性、可逆性和度量兼容传输的作用。通过实验,我们发现不同规模(从124M到8B)和结构(编码器/解码器)的训练Transformer表现出与我们理论一致的几何结构:有效注意力图在更深层收敛到稳定的几何算子,学习的传输自组织成近似缩放等距,且两种现象都随规模持续增强。总体而言,本文提供了一种将自注意力与经典几何算子联系起来的精确连接游走形式,以及一组从几何角度分析Transformer模型的算子级工具。

英文摘要

Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over the token-position graph and identify attention as a connection walk: messages are aggregated by a nonnegative walk matrix while being transported along each edge by a learned linear map. Within this framework, we prove that single-head attention (SHA) is exactly a connection propagation step with constant transport, and that multi-head attention (MHA) is exactly a single edge-dependent connection walk whose effective transport is an attention-gated mixture of headwise transports. We further clarify the conditions under which the corresponding generator reduces to a random-walk connection Laplacian, highlighting the roles of stochasticity, reversibility, and metric-compatible transports. Empirically, we find that trained Transformers across scales (from 124M to 8B) and structures (encoder/decoder) exhibit geometric structure consistent with our theory: effective attention graphs converge to stable geometric operators in deeper layers, learned transports self-organize into approximate scaled isometries, and both phenomena strengthen consistently with scale. Overall, the paper provides a precise connection-walk formalism that links self-attention to classical geometric operators, along with a set of operator-level tools for analyzing transformer models from a geometric perspective.

Comments29 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑