arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有注意力都平等:EEI权衡的定量综述

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

Aditya Singh

arXiv 2608.15459首次发表:更新:

AI 中文总结

本综述围绕注意力机制的效率-表达性-可解释性权衡,定量对比21种方法,梳理其多领域发展脉络,分析研究缺口并指明未来方向,支持粗粒度层级对比。

AI 中文摘要

注意力机制已推动机器学习发展十年,从神经机器翻译到具备通用推理能力的语言模型。本综述涵盖四个关联方向:面向序列到序列任务的注意力机制构建、向计算机视觉的适配、解决二次瓶颈的效率创新,以及可解释性进展。我们定义效率、表达性、可解释性三项标准,采用EEI评分框架对比21种方法,评分来自单一评估者,假设存在±1分的扰动范围。基于20万个样本的确定性蒙特卡洛分析显示,在该扰动模型下,平均67%-70%的样本会出现超过1个名次的变化;名次匹配的零模型复现了相似的稳定性特征,因此结果支持粗粒度层级对比而非细粒度排名。本综述梳理了从Bahdanau-Luong对齐到Transformer再到视觉架构的注意力发展脉络,涵盖固定与学习型稀疏注意力、线性注意力、含FlashAttention的IO感知精确算法、含Mamba的状态空间替代方案,还涉及诱导头、叠加态及注意力-SSM对偶性。此外,本综述提供结构化叙事综述、带有跨研究注意事项的基准综合、五个问题的研究缺口分析,以及2015-2026年演变时间线。结论将注意力研究定位为效率-表达性-可解释性前沿的拓展,并指出未来方向包括统一效率基准、混合架构的学习型路由、长度泛化及可扩展的机械可解释性。

英文摘要

Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose reasoning. This survey covers four connected threads: their formulation for sequence-to-sequence tasks, adaptation to computer vision, efficiency innovations that address the quadratic bottleneck, and advances in interpretability. We define three criteria: efficiency, expressiveness, and interpretability, and compare twenty-one methods using an EEI scoring framework. Scores come from a single rater with an assumed +/-1-point perturbation range. A deterministic Monte Carlo analysis with 200,000 samples shows that, under this perturbation model, rank changes of more than one position occur in 67-70% of samples on average. A rank-matched null model reproduces a similar stability profile, so the results support coarse tier-level comparisons rather than fine-grained rankings. The survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures. It reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba. It also covers induction heads, superposition, and the attention-SSM duality. We further provide a structured narrative review, a benchmark synthesis with cross-study caveats, a five-problem research gap analysis, and a 2015-2026 evolution timeline. We conclude by framing attention research as an expansion of the efficiency-expressiveness-interpretability frontier and identifying future directions including unified efficiency benchmarks, learned routing for hybrid architectures, length generalization, and scalable mechanistic interpretability.

Comments53 pages, 8 figures, 16 tables. Code and analysis artifacts: https://github.com/Nixon-H/not-all-attention-is-equal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑