发表机构
George Washington University; Cairo University; Carnegie Mellon University(乔治华盛顿大学; 开罗大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WNet用离散小波变换替代自注意力实现高效令牌混合,通过多种无注意力混合器降低长序列编码成本,在GLUE上验证了性能与速度优势。
AI 中文摘要
在Transformer中,令牌混合是让每个令牌从其他令牌获取信息的步骤,它主导了编码长序列的成本。自注意力机制能很好地完成这种混合:每个令牌按内容加权所有其他令牌,从而提供强大的上下文建模。但这种全对比较也是其成本随序列长度呈二次方增长的原因。我们引入WNet,一种Transformer编码器,用基于离散小波变换(DWT)的令牌混合取代自注意力。三种无注意力混合器通过不同方式重组尺度:线性融合、学习门控,或让每个令牌自行选择其尺度。一种混合变体仅在最后一层添加自注意力。感受野分析表明,由两抽头滤波器(如Haar)构建的小波混合器,无论网络多深,即使滤波器是可学习的,也从不关联固定块外的令牌。更长的滤波器在两层内即可覆盖整个序列。我们在C4的固定令牌子集上使用掩码语言建模预训练每个模型,并在GLUE上进行微调,采用一个受控设置,包括尺寸匹配的BERT和FNet基线以及一个无法混合令牌的对照组。令牌门控混合器在256个令牌时训练速度与注意力相当,在4096个令牌时快2.7倍。
英文摘要
In a Transformer, token mixing is the step that lets each token draw information from other tokens, and it dominates the cost of encoding long sequences. Self-attention does this mixing very well: every token weighs every other token by content, which gives strong contextual modeling. That all-pairs comparison is also why its cost grows quadratically with sequence length. We introduce WNet, a Transformer encoder that replaces self-attention with token mixing based on the discrete wavelet transform (DWT). Three attention-free mixers recombine the scales: by linear fusion, by learned gating, or by letting each token choose its scales. A hybrid adds self-attention in the last layer only. A receptive-field analysis shows that wavelet mixers built from two-tap filters, such as Haar, never relate tokens outside fixed blocks, however deep the network, even when the filters are learned. Longer filters reach the whole sequence within two layers. We pre-train every model with masked language modeling on a fixed-token subset of C4 and fine-tune on GLUE, using one controlled setup with size-matched BERT and FNet baselines and a control that cannot mix tokens. The token-gated mixer trains as fast as attention at 256 tokens and 2.7 times faster at 4,096.