双注意力残差
Dual Attention Residuals
浏览论文内容
中文总结 AI 辅助
研究提出双注意力残差(DAR),通过相互跨流寻址将多流交互引入历史检索,在多个模型中优于标准残差Transformer和注意力残差,路由消融等分析表明其能保留深度多样性,避免冗余和功能不平衡。
中文摘要 AI 辅助
近期工作沿两个互补轴扩展了Transformer残差路径:历史检索从较早深度选择信息,而多流方法维持多条残差轨迹。这些能力大多被孤立研究,为每个流分配独立检索器仍会阻止一条轨迹影响另一条轨迹的深度选择。我们提出双注意力残差(DAR),通过相互跨流寻址将多流交互引入历史检索。对于每个目标流,DAR从相反流的归一化状态计算深度权重并应用于目标流自身历史的值。检索到的状态组合用于不变的Transformer分支并通过受限门控写入更新;一种块形式变体对块级历史进行操作以控制开销。在参数从0.1B到1B的密集模型和一个7B稀疏MoE模型中,DAR始终优于标准残差Transformer和注意力残差。路由消融表明增益不能仅由额外流或值投影解释。表示和干预分析进一步表明相互跨流选择保留了深度多样性并避免了在替代双流设计中观察到的冗余或功能不平衡。
英文摘要
Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain multiple residual trajectories. These capabilities have largely been studied in isolation, and assigning an independent retriever to each stream still prevents one trajectory from influencing depth selection in another. We propose Dual Attention Residuals (DAR), which brings multi-stream interaction into historical retrieval through reciprocal cross-stream addressing. For each target stream, DAR computes depth weights from normalized states in the opposite stream and applies them to values from the target stream's own history. The retrieved states are combined for an unchanged Transformer branch and updated through constrained gated writes; a block-form variant operates on block-level histories to control overhead. Across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model, DAR consistently improves validation loss over standard residual Transformers and Attention Residuals. Routing ablations show that the gain cannot be explained by an additional stream or value projection alone. Representation and intervention analyses further show that reciprocal cross-stream selection preserves depth-wise diversity and avoids the redundancy or functional imbalance observed in alternative two-stream designs.