AI 中文总结
Extender 是一种日志结构化 Transformer,通过拼接通道减少注意力内存占用,在长上下文任务上超越标准 Transformer。
AI 中文摘要
我们提出了 Extender,它是标准 Transformer 架构的一种日志结构化变体。在标准 Transformer 中,每一层仅通过残差 $\mathbf{h}$(一个叠加通道)与后续层通信。Extender 增加了一个拼接通道 $\mathbf{x}$:每一层 $\ell$ 既发出一个残差更新 $\delta_\ell$(添加到 $\mathbf{h}$ 中),又发出一个更小的扩展 $\epsilon_\ell$(追加到 $\mathbf{x}$ 中)。虽然 FFN 和 $\mathbf{q}$ 都看到 $\mathbf{h}$,但注意力 $\mathbf{kv}$ 投影仅以 $\mathbf{x}$ 作为输入。因此,完全扩展的 $\mathbf{x}$ 包含所有层 $\mathbf{kv}$ 投影的完整输入,将持久注意力内存占用从 $2Ld_{model}$ 减少到 $\sum|\epsilon_\ell|$。我们发现,当 $|\epsilon_\ell|=32$ 时,在 199M-924M 参数规模下,Extender 在短上下文(CORE)任务上与 Transformer 的准确率相当;在 924M 参数规模下,在长上下文(RULER)任务上超过 Transformer 的准确率。对于我们的 1664 宽、924M 模型,Extender 的持久注意力内存占用比 MHA 小 $104\times$。内存节省随模型宽度增加而增长。
英文摘要
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $δ_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $ε_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|ε_\ell|$. We find that with $|ε_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.