三个令牌在非负核注意力中强制指数特征秩
Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
- CENIA(智利环境与可持续发展研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究对比全注意力与核注意力的差异,证明三令牌序列下非负核注意力需指数级特征,而密集softmax仅需线性特征,还推导了有限字母表多头草图模型的转录本下界。
AI中文摘要:
全注意力会暴露每一对令牌,而核注意力会将序列压缩为固定维度的草图。我们表明,当上下文长度达到包含两个竞争候选者的情况时,这种差异会呈现指数级增长。在布尔输入上的Min-IP(内积)问题中,秩一归一化核注意力能精确求解长度最多为2的所有序列。相比之下,任何在所有三令牌序列上成功且误差严格低于1/2的单归一化非负核注意力头,即使使用任意有限维度的逐令牌值和任意依赖查询的仿射读出,也需要2^Ω(m)个特征;而密集softmax仅需m维分数和常数温度即可完成相同任务。该结论在依赖位置的令牌映射和因果最终查询的情况下依然成立。随着上下文长度增加,该下界趋近于使用精确2^m个特征的实现。此外,对于确定性多头多层草图模型,其跨令牌通道具有有限字母表,我们证明了转录本下界与独立答案的数量呈线性关系,且与它们的字母表大小呈对数关系。
英文摘要:
How much feature rank does comparison require in kernel attention? On Min-IP over $m$-bit tokens, rank one solves every sequence of length at most two exactly. At length three, the minimum feature rank of one normalized nonnegative kernel-attention head is $2^{Θ(m)}$ for error strictly below $1/2$ on every input, even with arbitrary finite-dimensional tokenwise values and query-dependent affine readouts. Dense softmax solves this three-token task with $m$-dimensional scores and temperature constant in $m$. For every fixed number of heads $H$, the minimum total feature rank is $2^{Θ_H(m)}$ for the same error guarantee at exact length $H+2$ in one attention layer with affine mixing. These bounds also hold with position-dependent maps and a final causal query. For one head with polynomial readout of fixed degree at most $D$, rank one suffices at exact length $D+1$, while exact length $D+2$ requires exponential feature rank. With unrestricted exact-real decoding, a scalar rank-one construction solves the task at every finite length. This motivates a separate bound on total communication for deterministic models with finite-alphabet cross-token channels and any number of heads and layers. In this setting, correctness up to length $n$ requires $Ω(n\log m)$ bits over a range of lengths that grows exponentially with $m$.