发表机构
National Technical University of Athens; University of Bern; Amazon AGI; Archimedes RU, Athena RC(雅典国立技术大学; 伯尔尼大学; 亚马逊AGI; Archimedes研究单位,Athena研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出注意力感知路由(AAR),利用注意力权重滑动窗口的时频特征增强MoE路由器,仅训练路由参数,在GSM8K上提升3.37个百分点,并揭示路由与注意力的耦合回路及深度敏感性。
AI 中文摘要
在混合专家语言模型中,路由器通常基于token的隐藏状态来选择专家并分配权重,这仅利用了有限的上下文信息。我们提出了注意力感知路由(AAR),该方法利用从注意力权重滑动窗口中提取的时域和频域特征来增强路由器,这些特征代表了模型上下文状态的摘要,且与隐藏状态解耦。在保持基础Transformer完全冻结的情况下,我们仅训练路由参数,将路由作为唯一变量进行隔离。AAR在OLMoE上相比仅路由的SFT基线,在GSM8K上提升了+3.37个百分点。除了性能提升,我们还展示了路由和注意力形成了一个耦合回路:第l层的路由变化通过残差流传播,放大第l+1层的注意力汇聚效应,从而在不直接更新注意力机制本身的情况下重塑注意力。此外,AAR减少了长距离发散生成,错误答案变得更短,而正确答案的长度保持不变。最后,AAR对深度高度敏感:跨层不加区分地应用AAR可能会降低事实检索能力,而当其在网络更深层引入时,数学推理能力的提升得以保持。这种敏感性揭示了深度上的检索-推理张力,并使分层选择的AAR成为探测不同层注意力所携带的路由相关信息的一种受控探针。
英文摘要
In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model's contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT baseline on OLMoE. Beyond performance, we show that routing and attention form a coupled circuit: routing changes at layer l propagate through the residual stream to amplify attention sinks at layer l+1, reshaping attention without any direct update to the attention mechanism itself. Further, AAR reduces long diverging generation, with incorrect answers getting shorter, while correct answers remain unchanged in length. Finally, AAR is strongly depth-sensitive: applying it indiscriminately across layers can degrade factual retrieval, whereas mathematical reasoning gains persist when it is introduced deeper in the network. This sensitivity exposes a retrieval--reasoning tension across depth and makes layer-selective AAR a controlled probe of the routing-relevant information carried by attention at different layers.