arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单层注意力中布尔函数的头复杂度

The Head Complexity of Boolean Functions in Single-Layer Attention

Rajmohan Rajaraman, Ravi Sundaram, Amanuel Tesfaye

arXiv 2609.04046首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究定义了仅含一层注意力的模型中计算布尔函数所需的最小注意力头数(头复杂度),建立了其层级、下界、紧致性界及一般二元函数的通用界,明确了头数与维度、精度的不可替代性。

AI 中文摘要

单层自注意力能计算什么?我们研究头复杂度:仅含一层注意力的模型中,计算某个函数所需的最小注意力头数量。我们在该度量下建立了精确层级:k 个头可计算 k 比特奇偶校验,但无法计算 k+1 比特奇偶校验。该下界在Transformer可能利用的两种资源上是无条件的,在嵌入维度和数值精度无界时依然成立。证明基于交替和障碍:清除softmax分母后,所得决策多项式中的每个单项式至少遗漏k+1个输入比特中的一个,迫使其与奇偶校验的相关性消失。同一障碍为相关任务(包括研究充分的多跳诱导头任务)产生了下界。我们还建立了嵌入维度和数值精度的紧致性界:紧致性定理表明,任何可计算的函数都可通过由任务离散数据(即头数、字母表大小和长度)界定的嵌入维度和精度来计算,因此潜在无界的维度或精度无法替代头。最后,我们推导了一般二元函数的近乎匹配的通用界:2ⁿ个头足以计算每个n比特二元函数,对应其多线性展开中的每个单项式用一个头;计数论证显示,几乎所有此类函数都需要Ω(2ⁿ/n²)个头,即使维度和精度无界,该下界与上界的差距也仅为poly(n)因子。这些结果共同刻画了该模型中布尔计算的头需求。

英文摘要

What can a single layer of self-attention compute? We study head complexity: the minimum number of attention heads required to compute a function in a one-layer attention-only model. We establish an exact hierarchy under this measure: $k$ heads compute $k$-bit parity but cannot compute $(k+1)$-bit parity. The lower bound is unconditional in the two resources a transformer might otherwise exploit; it holds at unbounded embedding dimension and unbounded numerical precision. The proof rests on an alternating-sum obstruction: after clearing the softmax denominators, every monomial in the resulting decision polynomial omits at least one of the $k+1$ input bits, forcing its correlation with parity to vanish. The same obstruction yields lower bounds for related tasks, including the well-studied multi-hop induction-head task. We also establish compactness bounds for embedding dimension and numerical precision. Specifically, a compactness theorem shows that any function computable at all can be computed with embedding dimension and precision bounded by the discrete data of the task, namely, head count, alphabet size, and length. Thus, potentially unbounded dimension or precision provably cannot substitute for heads. Finally, we derive nearly matching universal bounds for general binary functions: $2^n$ heads suffice to compute every $n$-bit binary function, with one head per monomial in its multilinear expansion, while a counting argument shows almost all such functions require $Ω(2^n/n^2)$ heads. This lower bound matches the upper bound to within a $\operatorname{poly}(n)$ factor, even when dimension and precision are unbounded. Together, these results characterize head requirements for Boolean computation in this model.

Comments32 pages, 0 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑