arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09694cs.LGcs.CL

低秩注意力残差

Low-Rank Attention Residuals

Jonathan Su

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出低秩注意力残差(LR-AttnRes),包括投影式和切片式,通过使用低维键进行路由,解耦路由与残差内容,减少运算量并提升性能,证明深度路由在较少维度下也有效,还发布了代码和模型。

中文摘要 AI 辅助

注意力残差在大语言模型中用对先前子层输出的深度注意力取代固定残差和,但将每个输出用作全维键和值。这将路由与表示耦合,使深度路由分数随隐藏宽度\(d\)缩放。我们提出低秩注意力残差(LR-AttnRes),在使用\(r\)维键(\(r\ll d\))进行路由时保持全维残差值。投影式LR-AttnRes从现有输出投影中发射学习到的低秩键,解耦路由与残差内容,并在测试变体中实现最佳验证损失。切片式LR-AttnRes使用每个值的最后\(r\)维作为路由键,去除辅助键投影路径并减少残差侧浮点运算量,同时仍提高性能。全面扫描表明深度路由在维度远少于模型宽度时也有效。我们发布代码和模型以促进未来研究。

英文摘要

Attention Residuals (AttnRes) replace the fixed residual sum with depth-wise attention over previous sub-layer outputs in Large Language Models (LLMs), but use each output as both a full-dimensional key and value. This couples routing with representation and makes the cost of computing depth-routing scores scale with hidden width $d$. We propose Low-Rank Attention Residuals (LR-AttnRes), which keep full-dimensional residual values while using $r$-dimensional keys, with $r < d$, for routing. LR-AttnRes uses the last $r$ dimensions of each value as the routing key, reducing total residual-side FLOPs while still improving performance. Comprehensive sweeps across the number of blocks ($N$) and $r$ show that depth-wise routing can be effective with far fewer dimensions than the model width. At both $1$B and $4$B parameters with $r = d/4$, LR-AttnRes achieves lower final validation loss, higher average downstream accuracy, and higher measured training-step throughput than standard AttnRes. We also provide a fused kernel supporting standard and low-rank routing. We release all code, the kernel, and all trained models to facilitate future research.

↑