学习查询和键即是一切:用结构化变换替代值投影
Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms
浏览论文内容
中文总结 AI 辅助
本文提出双头Transformer,用结构化变换(如WHT、DCT等)替代值投影以减少参数和缓存,在ImageNet上优于三头结构。
中文摘要 AI 辅助
为了减少Transformer的参数数量和缓存内存需求,我们引入了双头Transformer以替代三头结构。我们研究了Walsh-Hadamard变换(WHT)、离散余弦变换(DCT)、离散傅里叶变换、基于滤波器组的Shearlet变换以及乘法避免(MA)算子来构建双头。我们将空间块及其正交变换(或Shearlet和MA算子)以类似于注意力块的结构相结合。在ImageNet上,我们获得了优于三头Transformer的结果。文中展示了大量的仿真示例。
英文摘要
To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.
发表机构
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。