arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18145cs.LG

无需搜索即可到达每个位置:超立方体上的旋转稀疏布线作为注意力的替代方案

Reaching Every Position Without Searching: Rotating Sparse Wiring on the Hypercube as a Substitute for Attention

发表机构早稻田大学
查看机构详情
  • Waseda University(早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

Yoshiaki Takashita

首次发表
浏览论文内容

中文总结 AI 辅助

提出超立方体上逐层旋转的固定稀疏布线替代注意力,以对数层数实现全位置信息传播,在语言建模中达到更低损失且大幅减少计算和参数。

中文摘要 AI 辅助

注意力机制在每一层、对每个输入,都要付出搜索连接对象的代价。我们探究使用固定、稀疏且逐层简单旋转的布线能达到何种效果。将序列的$n$个位置视为$\nlog_2 n$维超立方体的顶点,并在第$\ell$层将每个位置连接到其沿维度$\ell \bmod \log_2 n$的邻居,来自每个位置的信息在$\log_2 n$层内即可到达所有其他位置,每层仅需$2n$条连接,而非$n^2$条。在一个除非所有位置都被到达否则无法解决的合成任务上,这种旋转方式以$1/32$的连接数匹配了全连接布线,而相同稀疏模式在各层固定不变时则失败;关键在于每个维度都被触及,而非顺序。在公共语料库(enwik8的前$12$M个字符)的字符级语言建模中,一个在十六个稀疏层中保留两个注意力层的混合模型,在相同宽度和相同步数预算下(每种三个种子,无重叠),比全注意力模型的保留损失低$0.06$比特每字符,同时连接数仅为后者的$1/7$,参数减少$42\\%$,墙钟时间减少$2.4$倍;纯旋转调度与混合模型持平。在第二个混合日语、英语和代码的语料库上,排序相同,差距扩大至$0.16$。在两者上,可用的学习率窗口是注意力的四到八倍。我们还报告了无效的方法——学习坐标,以及一种“动力学”变体,其表面上的收益实为饱和核的伪影——以及我们发现在此规模下得出任何结论所必需的测量纪律(冻结语料库、全覆盖评估、以种子散布作为排名标准)。

英文摘要

Attention pays, at every layer and for every input, the cost of searching for whom to connect. We ask how far one can get with wiring that is fixed, sparse, and simply rotated from layer to layer. Treating the $n$ positions of a sequence as the vertices of a $\log_2 n$-dimensional hypercube and connecting each position, at layer $\ell$, to its neighbour along dimension $\ell \bmod \log_2 n$, information from every position reaches every other in $\log_2 n$ layers with $2n$ links per layer instead of $n^2$. On a synthetic task that is unsolvable unless all positions are reached, this rotation matches all-to-all wiring at $1/32$ of the links, while the same sparse pattern held fixed across layers fails; what matters is that every dimension is touched, not the order. On character-level language modelling of a public corpus (the first $12$M characters of enwik8), a hybrid that keeps two attention layers among sixteen sparse ones reaches $0.06$ bits-per-character lower held-out loss than a fully attentive model of the same width at the same step budget (three seeds each, no overlap), with $1/7$ of the links, $42\%$ fewer parameters, and $2.4\times$ less wall-clock time; the purely rotated schedule is level with the hybrid. The same ordering holds on a second corpus of mixed Japanese, English and code, where the gap widens to $0.16$. The usable learning-rate window is four to eight times wider than attention's on both. We also report what did not work - learned coordinates, and a "dynamics" variant whose apparent gains turned out to be an artefact of a saturated kernel - and the measurement discipline (frozen corpus, full-coverage evaluation, seed spread as the bar for ranking) that we found necessary to say anything at all at this scale.

补充信息

↑