Softmax注意力的近似秩:尖锐几何定律与鲁棒交互维度
Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention
浏览论文内容
中文总结 AI 辅助
该研究明确了控制Softmax注意力秩复杂度的几何规律,推导了球形、全球等支撑下的近似秩公式,经BERT-base校准集验证,将支撑几何与交互几何的作用区分开。
中文摘要 AI 辅助
何种几何结构控制归一化Softmax注意力的秩复杂度?我们研究最大行ℓ₁近似秩,即保留所有有界向量值输出的最小无约束秩。两项尖锐最坏情况定律分离出支撑几何:对于固定的d和误差ε,球形自注意力的秩为Θ_{d,ε}(min{n,(1+β)^((d-1)/2)});而全球几何增加一个径向自由度,当β≥β₀(d,ε)且n≥C_d e^{β/8}时,秩为Θ_{d,ε}(β^{d/2})。对于固定头,行Softmax消去行标量logit方向:剩余可见的查询-键交互维度r产生每实例上界定律r/2,且有界构造表明该指数是极小极大尖锐的。近似交互子空间产生显式残差输出误差和容差索引的SVD维度。在84头BERT-base校准集上,我们观察到在多头-温度设置下有效维度适度降低,且与有限构造秩上界证书呈正相关。这些结果共同将设定最坏情况温度缩放的支撑几何,与控制每头近似复杂度的Softmax可见交互几何区分开。
英文摘要
How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\ell_1\) approximation rank \(r_\varepsilon(A)\), exactly the least rank achieving uniform error over all bounded vector-valued values. Row softmax exposes the intrinsic interaction \(C=P_m(\log A)P_N\), whereas invertible \(Q/K\) gauges leave \(A\) fixed while changing the Euclidean geometry of a chosen query/key factorization. We replace that coordinate-dependent description by a projective residual \(q(C-T)\) and an attained factor-radius size \(κ(T)\). For every rank-\(r\) retained interaction with \(τ(T)<\varepsilon\), we prove $$ r_\varepsilon(A)\le \min\left\{ N,\; C_r\left( 1+\frac{κ(T)} {(\varepsilon-τ(T))^2} \right)^{r/2} \right\}, $$ with the same unknown dimension constant as the underlying weighted Gibbs-row cover. The profile is gauge invariant, termwise no worse than native retained-subspace bounds at the same declared dimension, and has a worst-case sharp \(r/2\) size exponent at fixed \(r\) and \(\varepsilon\). We then measure \(r_\varepsilon(A)\) directly on learned attention using 9,978 certified brackets across BERT, GPT-2, Qwen2.5, and two ViT checkpoints; where certificates do not close, the optimum remains interval-valued. A pre-specified 2,302-cell held-out study further shows that the historical native-coordinate geometry block contains coarse, mostly head-level information but no detectable incremental information beyond a strong calibrated baseline. The new intrinsic descriptor is not evaluated in that study. Together, the theory and measurements distinguish an operator-intrinsic complexity control from a stronger empirical explanation that the learned-head evidence does not support.
发表机构
- Nanyang Technological University(南洋理工大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。