AI 中文总结
该研究提出混合门控注意力(HyGA)框架,通过三类门控策略等提升注意力机制的有效性与效率,实验显示其在训练损失、下游性能及不同计算成本下均优于门控注意力。
AI 中文摘要
门控注意力是缓解注意力下沉问题并提升注意力表征能力的有效方法。为进一步拓展其在有效性-效率帕累托前沿的表现,我们提出包含三类门控策略的混合门控注意力(HyGA)框架。具体而言,这些门控利用注意力多个阶段的不同信息,从多个视角协同构建元素级/头级门控,捕捉头内及头间信息交互。通过混合门控组件,HyGA可提供多源调制信号,实现对信息流更全面的控制并提升注意力表征能力。我们还引入低秩矩阵分解与可学习注意力下沉,进一步提升训练效率与稳定性。实验中,我们基于不同骨干网络在广泛使用的基准上评估HyGA,结果显示与门控注意力相比,HyGA在训练损失及各类下游任务性能上均有全面提升,且在不同计算成本下均达到最优性能,相关全面模型分析也有助于更好地理解该方法。所提出的HyGA为构建更有效、高效且稳定的注意力机制提供了思路。
英文摘要
Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of gating strategies. Specifically, these gates leverage diverse information from multiple stages of attention, and collaboratively build element-wise/head-wise gating from multiple perspectives, capturing intra-head and cross-head information interactions. Through our hybrid gating components, HyGA could provide multi-source modulation signals, enabling more comprehensive control over information flow and improving the representational capacity of attention. We also introduce low-rank matrix decomposition and learnable attention sink to further enhance training efficiency and stability. In experiments, we evaluate HyGA on widely-used benchmarks based on different backbones. The experimental results show that our HyGA comprehensively improves both training loss and various downstream performances compared with Gated attention. HyGA has also been verified to achieve the best performance at different computation costs, with comprehensive model analyses for better understanding. The proposed HyGA sheds light on a more effective, efficient, and stable attention mechanism.