发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出代数形式化方法推导因果掩码有限精度变换器的表达能力,结合数值语义与半群分类,明确不同注意力机制对应的记忆类型及表达边界,且边界均为紧的。
AI 中文摘要
因果掩码有限精度变换器可求解任意长度输入下的何种决策问题?现有答案常依赖理想算术,但有限精度下,舍入与求值顺序会改变注意力保留的信息,进而影响模型可计算内容。我们开发了一种代数形式化方法,直接从模型的实现动态推导其表达能力,核心对象是其记忆:注意力计算出的有限内部状态,总结前缀中所有未来查询可用的信息。每层中,每个注意力头独立更新自身状态,层间则分层组合,为从模型假设到表达能力边界提供了统一路径。将此方法应用于无位置嵌入的变换器,我们得到由特定数值语义下注意力类型决定的表达能力层级:宽度1的滑动窗口注意力支持有界后缀记忆,改进的软注意力支持不可逆的清单式状态,两者结合可产生两种机制的相互作用,普通从左到右浮点软注意力则能实现比上述任何一种更具表达力的记忆操作。代数上,这四种情况分别对应确定半群、R-平凡半群、局部R-平凡半群和非周期半群。在显式自由连接假设下,所有四个边界均为紧边界。
英文摘要
What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains and therefore what the model can compute. We develop an algebraic formalization that derives expressivity directly from the model's implemented dynamics. Its central object is its memory; the finite internal state computed by attention that summarizes the information from the prefix available to all future queries. Each attention head updates its own state independently within a layer, while layers compose hierarchically, providing a uniform route from model assumptions to expressivity bounds. Applying this method to transformers without positional embeddings, we obtain an expressivity hierarchy governed by the attention type under specific numerical semantics. Width-one sliding-window attention supports bounded-suffix memory, while a modified form of soft attention supports irreversible, checklist-like state, and combining the two mechanisms provides an interplay of both. Ordinary left-to-right floating-point soft attention can realize more expressive memory operations than any of the above. Algebraically, the four cases correspond to definite, R-trivial, locally R-trivial, and aperiodic semigroups. Under an explicit free-wiring assumption, all four bounds are tight.
Comments12 pages, 2 figures