发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明仅用两个冻结的高斯初始化单头残差softmax注意力块,即可在连续和有限深度下实现任意序列集合间的通用插值,并刻画了因果掩码的限制。
AI 中文摘要
通用逼近是学习架构从缩放定律中获益所必需的一个定性属性。虽然它在各种神经架构和随机特征模型上普遍得到验证,但通常涉及无限宽度极限。在这项工作中,我们专注于深度自注意力模型,并考虑相反的“对偶”机制,其中逼近能力完全由深度提供,并具有跨层的强参数共享,这受到诸如循环变换器(Looped Transformers)等近期模型的启发。更具体地说,我们询问是否可以找到一组预定义的有限参数,每个参数定义一个注意力块,使得由此产生的有限变换集合能够将任意 $N$ 个长度为 $n$ 的 token 序列的集合映射到任意其他 $N$ 个长度为 $n$ 的 token 序列的集合。关键在于,这些变换是独立于输入和输出集合而固定的:只有块应用的顺序、它们的符号以及它们的持续时间取决于特定的插值任务。我们的主要结果确立了这一点,对于残差 softmax 注意力,仅使用两个冻结的单头块,其投影矩阵采用高斯初始化。该结果在连续深度和有限深度下均成立。我们还刻画了因果掩码所施加的限制,并建立了相应的通用插值保证。
英文摘要
Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of $N$ sequences of $n$ tokens to any other collection of $N$ sequences of $n$ tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.