注意力机制与状态空间模型在上下文学习中的比较分析
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
浏览论文内容
中文总结 AI 辅助
本文提出信念几何框架,统一比较注意力与状态空间模型在上下文学习中的能力,发现SSM在信念维持和位置组装上有优势,而softmax注意力在内容寻址上具有指数级宽度优势,并通过实验验证了这些架构见解的普适性。
中文摘要 AI 辅助
Transformer 和状态空间模型(SSM)是两种突出的序列学习架构,然而对它们的比较在很大程度上仍是经验性的,且现有的理论分析通常针对特定任务或受架构限制。在本文中,我们开发了信念几何(belief geometry),一个用于比较广泛类别的注意力机制和 SSM 的表征能力的统一分析框架。从上下文线性回归的广义公式出发,并以累积贝叶斯遗憾(cumulative Bayes regret)作为度量,我们在信念几何中抽象出许多序列学习问题所需的三种能力:证据组装(evidence assembly)、信念维持(belief maintenance)和寻址(addressing)。然后,我们研究了公式中隔离这些能力的三种情况,并得出尖锐的架构教训:对于信念维持,SSM 在平稳聚合核上达到最优遗憾;对于位置组装,SSM 具有记忆优势;对于内容寻址,softmax 注意力相对于 sigmoid 选择的 SSM 具有指数级的宽度优势。使用 LLaMA 型 Transformer 和 Mamba-2 的实验表明,这些架构见解可扩展到我们可解析处理的类别和线性回归测试平台之外。
英文摘要
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.