arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24440cs.LG

通过稀疏自编码器比较状态空间模型与Transformer中的潜在概念形成

Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder, Ashwini M Joshi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过稀疏自编码器对比Mamba与Transformer的潜在表征,发现两者在语义理解上无系统性差异,仅少数特征在语法边缘处存在分歧。

中文摘要 AI 辅助

Transformer自注意力的二次方扩展性推动了亚二次方选择性状态空间模型(SSMs)如Mamba的采用,这些模型将过去的上下文压缩为固定大小的循环隐藏状态。这种严格的信息瓶颈为机制可解释性提出了一个基础性问题:SSMs和Transformer是否学习到根本不同的潜在表征?在本工作中,我们使用稀疏自编码器(SAEs)在1000万token语料库上对Mamba-130m和Pythia-70m进行了大规模、特征级别的对应分析。与预测广泛架构差异的假设相反,我们发现没有证据表明架构之间存在系统性的表征差异:在观察到的Jaccard分布中,99.98%的Mamba特征聚集在上对齐边界附近,为普遍性假说提供了初步的特征级别支持。我们进一步识别并定性描述了这一微小比例(0.02%)的差异特征,发现其模式与以下假设一致:循环瓶颈选择性地限制了对刚性语法的解析,而非广泛的语义本体。我们证明,虽然Pythia的无约束注意力允许对不同格式边缘情况进行单语义分解,但Mamba被迫将不相关的句法异常压缩到多语义的“杂物抽屉”神经元中,以保留状态容量。总的来说,这些结果表明架构路由机制可能对核心语义理解影响甚微,表征差异仅限于极端结构边缘。

英文摘要

The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs and Transformers learn fundamentally distinct latent representations? In this work, we employ Sparse Autoencoders (SAEs) to conduct a large-scale, feature-level correspondence analysis between Mamba-130m and Pythia-70m over a 10-million token corpus. Contrary to hypotheses predicting widespread architectural divergence, we find no evidence of systematic representational divergence between architectures: across the observed Jaccard distribution, 99.98% of Mamba features cluster toward the upper alignment boundary, providing preliminary feature-level support for the Universality Hypothesis. We further identify and qualitatively characterize this microscopic fraction (0.02%) of diverging features, finding patterns consistent with the hypothesis that the recurrent bottleneck selectively limits the parsing of rigid syntax rather than broad semantic ontology. We demonstrate that while Pythia's unconstrained attention permits the monosemantic decomposition of distinct formatting edge-cases, Mamba is forced to compress unrelated syntactical anomalies into polysemantic "junk drawer" neurons to preserve state capacity. Collectively, these results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.

发表机构

  • PES University(PES大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑