arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11197cs.LGcs.CL

超越特征袋:稀疏自编码器中的集合级不稳定性

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak

首次发表
浏览论文内容

中文总结 AI 辅助

该研究以稀疏自编码器(SAE)隐集合重叠度为相似度度量,发现SAE特征不通过简单特征袋语义组合,其激活集合与人类概念判断存在显著不匹配。

中文摘要 AI 辅助

Shani等人(2026)表明,大语言模型(LLM)表示大体上能恢复人类的类别边界,但无法反映细粒度的典型性结构。他们的分析采用了密集模型表示上的余弦相似度。我们使用活跃稀疏自编码器(SAE)隐集合的重叠度作为更具可解释性的相似度度量,重新审视其方法。我们首先验证该集合级度量的意义:SAE隐集合可在受控玩具模型中恢复类组合结构,并在自然文本中诱导语义连贯的邻域。将人类概念分析扩展到SAE集合相似度后,我们发现SAE激活集合相比密集嵌入或残差流状态,并未更忠实地恢复人类类别边界或类别内典型性,而是追踪模型内部的相似度结构。为进一步探究该差距,我们在受控语义修改下研究活跃隐集合,发现人类对概念变化的判断与SAE活跃集合的变化存在显著不匹配。我们将此解释为证据:在非理想环境中,SAE特征并非通过简单的特征袋语义组合。

英文摘要

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.

发表机构

  • Heidelberg University(海德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑