发表机构
Chemnitz University of Technology; University of Göttingen(开姆尼茨工业大学; 哥廷根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出由单变量样条和多项式指数核构建的加性核,在保持无限容量的同时实现拟线性计算,并在长序列上优于现代softmax后端。
AI 中文摘要
具有softmax注意力的Transformer的评估成本随序列长度呈二次方增长。核注意力通过将softmax替换为更一般的核函数来解决这一问题。本文旨在识别那些既能保留注意力表达力又能实现拟线性计算的核。为了量化表达力,我们为每个核引入一个容量,衡量注意力矩阵能近似单位矩阵的最大序列长度。因此,更高的容量表示更强的表达力。我们证明了诸如softmax、高斯和拉普拉斯等有表达力的核具有无限容量。相比之下,常见的拟线性核(例如由有限维特征映射导出的核)表现出有限容量。作为解决方案,我们提出了由单变量样条和多项式指数核构建的加性核。我们证明了这些核在通过排序实现拟线性计算的同时,保持了无限容量。最后,我们高效地实现了加性排序核,并将其与现代softmax后端进行基准测试,展示了在长序列上的优势。
英文摘要
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To quantify expressivity, we introduce a capacity for each kernel, measuring the maximum sequence length for which the attention matrix can approximate the identity. A higher capacity thus indicates greater expressivity. We show that expressive kernels like softmax, Gauss, and Laplace have infinite capacity. In contrast, common quasi linear kernels, such as those derived from finite dimensional feature maps, exhibit finite capacity. As a solution, we propose additive kernels constructed from univariate spline and polynomial exponential kernels. We prove that these maintain infinite capacity while allowing quasi linear computation via sorting. Finally, we implement additive sorting kernels efficiently and benchmark them against modern softmax backends, demonstrating advantages for long sequences.