发表机构
The Hebrew University of Jerusalem(耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过多尺度分析发现,尽管Transformer与SSM在全局表示几何上存在差异(前者主导方向突出,后者均匀分布),但在局部语义流形层面两者功能高度收敛,且有效容量与概念编码维度相似。
AI 中文摘要
近期诸如Mamba等状态空间模型(SSMs)尽管依赖根本不同的架构,却实现了与Transformer相当的语言建模性能。这引发了一个重要问题:这些结构差异如何影响其内部表示的几何形状和功能性质?我们通过对Transformer、SSM和混合架构中的表示进行多尺度分析来研究这一问题。首先,我们发现SSM将其表示信息均匀分布到所有维度,而Transformer的表示则严重受单一主方向主导。通过评估混合架构,我们观察到在每一注意力层之后,表示空间越来越偏向单一主导方向。接下来,我们通过可压缩性探索表示的不同几何分布如何影响表示容量。令人惊讶的是,我们发现尽管几何结构截然不同,两种架构的有效容量却紧密匹配。我们进一步研究这种偏斜的几何是否影响概念的编码方式。利用秩约束探针,我们证明两种架构在维度惊人相似的子空间中编码概念。此外,我们证明Transformer的主导主方向本身并不编码更多概念信息。最后,我们聚焦于流形之间的对齐,通过分析特定主题的表示或查看标记的最近邻域,发现它们高度对齐。最终,我们的分析表明,尽管Transformer和SSM引发了对潜在空间的不同使用,但它们在局部语义流形层面展现出显著的功能收敛。
英文摘要
Recent state-space models (SSMs) such as Mamba achieve language modeling performance comparable to transformers despite relying on fundamentally different architectures. This raises an important question: how do these structural differences influence the geometry and functional nature of their internal representations? We study this question through a multi-scale analysis of representations in transformers, SSMs, and hybrid architecture. First, we find that SSMs distribute their representational information evenly across all dimensions, whereas transformer representations are heavily dominated by a single principal direction. By evaluating hybrid architectures, we observe that the representation space becomes increasingly skewed toward a single dominant direction after each attention layer. Next, we explore how the different geometric spread of representations impacts representational capacity through compressibility. Surprisingly, we find that despite their contrasting geometric structures, both architectures exhibit tightly matched effective capacities. We further investigate whether this skewed geometry affects how concepts are encoded. Using rank-constrained probes, we demonstrate that both architectures encode concepts in subspaces of surprisingly similar dimensionality. Furthermore, we demonstrate that the transformers' dominant principal direction does not inherently encode more conceptual information. Finally, we zoom in and examine the alignment between manifolds, either by analyzing representations of specific topics or by looking at the nearest neighborhoods of tokens, and find that they are highly aligned. Ultimately, our analysis suggests that while transformers and SSMs induce different usage of latent space, they display a striking functional convergence at the level of local semantic manifolds.