发表机构
Johns Hopkins University; University of Cambridge(约翰霍普金斯大学; 剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用稠密图极限理论证明图变换器的注意力矩阵随规模增长收敛到稳定的注意力图极限,并提出估计与诊断方法,实验验证了其稳定性与可迁移性。
AI 中文摘要
图变换器为每个注意力头生成一个密集的$n\ imes n$矩阵,其中包含学习到的成对交互信息。我们提出一个基本问题:随着$n$的增长,这些由注意力诱导的图是否收敛到一个稳定的极限对象,还是学习到的交互模式仍然是无结构的且依赖于大小?我们利用稠密图极限理论来回答这个问题,将每个注意力矩阵视为来自底层核(即注意力图极限)的有限样本,并在切割距离下研究围绕该极限的集中性。我们推导出一个最坏情况下的方差界,该界不需要对图极限做任何假设,以及一个基于非参数估计理论的更尖锐的正则性感知界。为了使理论可操作化,我们提出了一种先规范化再分块平均的流程,用于估计数据集级别的注意力图极限,以及一个基于方差的诊断方法,用于测试注意力是否具有稳定的连续体描述。在多个图基准上的实验表明,在若干数据集上,学习到的注意力稳定到特定于数据集的图极限结构;经验切割距离和切割范数方差随$n$的减小与我们的界一致;并且注意力图极限可以迁移到更大的图大小,误差随$n$递减。
英文摘要
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-dependent? We answer this using dense graph limit theory, treating each attention matrix as a finite sample from an underlying kernel---an \emph{attention graphon}---and studying concentration around this limit under the cut-distance. We derive a worst-case variance bound requiring no assumptions on the graphon, and a sharper regularity-aware bound based on nonparametric estimation theory. To operationalize the theory, we propose a canonicalize-then-block-average pipeline for estimating dataset-level attention graphons, and a variance-based diagnostic for testing whether attention admits a stable continuum description. Experiments across multiple graph benchmarks show that learned attention stabilizes to dataset-specific graphon structure on several datasets; that empirical cut-distance and cut-norm variance decreases with $n$ consistent with our bounds; and that attention graphons transfer to larger graph sizes with error decreasing in $n$.
Comments42 pages, 32 figures