arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

角色解耦注意力残差:跨深度分离匹配与内容检索

Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth

Kehan Wang

arXiv 2608.01075首次发表:更新:

AI 中文总结

该研究提出角色解耦注意力残差(RD-AttnRes),将注意力的匹配与内容检索解耦,在多参数规模模型上验证其可提升语言建模性能,证实二者需从残差层级差异化读取。

AI 中文摘要

深度路由残差架构允许Transformer层检索更早的表示,而非仅继承前一层的状态。然而,现有的块注意力残差使用单一依赖内容的深度混合来构建查询、键和值的输入,该设计耦合了两种功能不同的决策:查询和键决定注意力的匹配位置,值决定要检索的内容。因此,我们提出疑问:匹配和内容检索是否必须从同一深度读取?我们引入角色解耦注意力残差(RD-AttnRes),这是一种极简扩展,在查询和键之间共享一条深度路由,同时在相同残差源上学习独立的值路由。将两条路由查询绑定可恢复原架构,而解耦仅在每层增加一个模型宽度向量,且不引入额外的token-to-token注意力操作。我们使用冻结配对预训练协议在FineWeb-Edu上评估RD-AttnRes,针对1.2亿参数和3.43亿参数模型各设置5个匹配种子,训练预算为20亿token。RD-AttnRes在全部10组匹配对比中均提升了验证集负对数似然,1.2亿参数模型的平均降幅为0.0301,对应困惑度降低2.97%;3.43亿参数模型的平均降幅为0.0247,对应困惑度降低2.43%。早期预算控制实验表明,额外参数数量、重复路由执行或固定值路由均无法复现该提升。路由诊断进一步揭示,查询-键与值的深度分布存在持续差异。这些结果表明,在评估的训练范式内,注意力匹配与内容检索受益于对残差层级的差异化读取。

英文摘要

Depth-routing residual architectures allow Transformer layers to retrieve earlier representations instead of inheriting only the immediately preceding state. Existing Block Attention Residuals, however, use a single content-dependent depth mixture to construct the inputs to queries, keys, and values. This design couples two functionally different decisions: queries and keys determine where attention matches, whereas values determine what content is retrieved. We therefore ask whether matching and content retrieval should be forced to read from the same depth. We introduce Role-Decoupled Attention Residuals (RD-AttnRes), a minimal extension that shares one depth route between queries and keys while learning an independent value route over the same residual sources. Tying the two routing queries exactly recovers the parent architecture, while decoupling them adds only one model-width vector per layer and introduces no additional token-to-token attention operation. We evaluate RD-AttnRes using a frozen, paired pretraining protocol on FineWeb-Edu with five matched seeds for both 120M- and 343M-parameter models and a 2.0B-token training budget. RD-AttnRes improves validation negative log-likelihood in all 10 matched comparisons. The mean reductions are 0.0301 and 0.0247, corresponding to perplexity reductions of 2.97 percent and 2.43 percent at 120M and 343M parameters, respectively. Early-budget controls indicate that neither the additional parameter count, duplicated routing execution, nor a fixed value route reproduces the improvement. Routing diagnostics further reveal persistent divergence between the query-key and value depth distributions. These results suggest that, within the evaluated training regime, attention matching and content retrieval benefit from distinct reads over the residual hierarchy.

Commentsa improvement of attnres

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑