arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

索引:起点与终点

Indexing: the Beginning and the End

Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia

arXiv 2607.22361首次发表:更新:

发表机构

CENIA; Pontifical Catholic University of Chile(CENIA研究所; 智利天主教大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究现代深度学习架构中信息瓶颈,通过索引原语视角,引入因果复杂度,分析索引在输入不同位置时各架构解决索引原语的能力,得出不可能性结果,实验与理论定性相符。

AI 中文摘要

我们通过索引原语的视角研究现代深度学习架构(循环神经网络、softmax 变换器、线性注意力变换器和状态空间模型)中的信息瓶颈。在此原语中,输入由 n 位和一个从 1 到 n 的整数 i(称为索引)组成,输出等于第 i 位的值。我们引入了掩码架构的因果复杂度。当索引出现在输入末尾时,具有低因果复杂度的架构无法在任何固定层数内解决索引原语问题,低参数循环神经网络、状态空间模型和掩码线性注意力变换器受此限制。而小型 softmax 变换器可在一层解决,非掩码线性注意力变换器可在两层解决。当索引出现在开头时,小型循环神经网络能在一层解决,其他架构则需两层。所有不可能性结果都是无条件的,实验也定性地与理论相符。

英文摘要

We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits and one integer $i$ from $1$ to $n$ called the index, and the output equals the value of the $i$-th bit. We introduce causal complexity for masked architectures. We show that architectures with low causal complexity cannot solve the indexing primitive in any constant number of layers when the index appears at the end of the input. In particular, this limitation applies to low-parameter RNNs, SSMs and masked linear-attention transformers. In contrast, small softmax transformers can solve it in one layer, while non-masked linear-attention transformers can solve it in 2, which separates them from their masked counterparts. In turn, when the index appears at the beginning, we show that small RNNs are capable of solving this task in 1 layer, while all the other architectures require 2. All our impossibility results are unconditional and apply even to models that employ infinite-precision real arithmetic. Moreover, experiments for up to $n=64$ qualitatively align with our theory: configurations with low-parameter theoretical solutions learn the indexing task easily, while configurations that do not admit such theoretical solutions struggle to learn as the sequence length grows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑