arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TANGO:用于自然语言与形式语言建模的令牌聚合门控非线性算子

TANGO: Treating Tokens as Operators

Joshua Nunley

arXiv 2608.22117首次发表:更新:

发表机构

Luddy School of Informatics, Computing, and Engineering; Indiana University Bloomington(勒迪信息学、计算与工程学院; 印第安纳大学伯明顿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出TANGO和WANGO两种模型,替换Transformer的自注意力与前馈子层,经对比实验,TANGO在多数据集上验证负对数似然最优,WANGO在线性复杂度架构中表现最佳且优于Recurrent Transformer++。

AI 中文摘要

标准Transformer块将自注意力中的跨令牌交互与在每个位置独立应用的非线性前馈网络分开。我们引入TANGO模型(令牌聚合门控非线性算子),将这两个子层替换为一个跨令牌门控残差更新。每个源令牌生成一个SwiGLU门控向量,查询-键相似度确定每个目标的源门控加权平均,所得门控重新缩放投影后的目标特征。TANGO为每个因果可见源分配单独权重,序列长度复杂度为二次。WANGO模型(非线性门控算子的窗口聚合)在近期窗口内保留相同未归一化分数,对更早源使用正特征图前缀统计量,在固定窗口和特征维度下序列长度复杂度为线性。我们将TANGO和WANGO与Recurrent Transformer++、全注意力GAU及FLASH对比,所有模型均有约4430万非嵌入参数,分三次匹配运行训练。TANGO、WANGO和Recurrent Transformer++将一个共享块应用四次,其余架构使用四个独立块。TANGO在FineWeb-Edu、Lean和DeepMind Mathematics上获得最低平均验证负对数似然,尽管其解析前向传递操作数最多;WANGO在序列长度复杂度为线性的架构中获得最低平均FineWeb-Edu负对数似然,且在几乎相同的解析前向传递乘积累加计数下优于Recurrent Transformer++。

英文摘要

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO computes a nonlinear feature-wise gate at each source token. Attention averages these gates for each destination. The average modulates a linear projection of the destination and forms the diagonal core of a source-conditioned linear operator. We test this proposal by comparing full-prefix and windowed TANGO with looped and untied Transformers, the Gated Attention Unit (GAU), and Fast Linear Attention with a Single Head (FLASH) on web text, Lean formal mathematics, DeepMind Mathematics, and code. The comparison uses two parameter scales, two depths, and three seeds. Checkpoints are selected on development data and evaluated on held-out test data. At matched parameters and training data, full-prefix TANGO has the lowest mean test negative log-likelihood in all 16 settings. In eight additional combinations of size and dataset, its development loss never increases as depth rises from 4 to 8 to 16, whereas the looped Transformer's loss increases in four. Full-prefix TANGO is computationally expensive because it averages wide gates over every visible source. To reduce this cost, we evaluate a variant with three narrower gated-projection sets assigned to the first, middle, and last applications. Across four FineWeb-Edu settings, this variant achieves 3.26 to 3.45 times the throughput of TANGO and 75% to 96% that of the looped Transformer. Its mean development negative log-likelihood is lower than TANGO's in three settings and 0.023 higher in the fourth, while remaining lower than both Transformer baselines in all four.

Comments23 pages, 1 figure, 15 numbered tables. Updated with depth-16 results and a faster three-projection-set TANGO variant

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑