arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29362cs.CL

词性作为SAE潜在空间中的涌现类别

Parts-of-Speech as Emergent Categories in SAE Latent Space

Alessandro Bondielli, Lucia Passaro, Serena Auriemma, Alessandro Lenci

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过词性分类测试,发现SAE潜在空间以分布式特征组而非原子特征编码形态句法信息,且类别可恢复性不依赖词汇记忆。

中文摘要 AI 辅助

稀疏自编码器(SAEs)为检查语言模型表示提供了一种有前景的方式,但其潜在特征所暴露的语言结构类型仍不清楚。我们使用词性(PoS)类别作为受控测试案例,研究形态句法信息是由单个潜在特征编码,还是由结构化特征组编码。我们发现,词性区分可以从SAE激活中高度恢复,但并非与一对一潜在特征/类别映射对齐。这种可恢复性不能归结为词汇记忆,且开放和封闭词性类别差异显著。类别由稀疏潜在特征的紧凑组支持,各标签间存在显著差异。这些组在保留数据上保持稳定,同时相关类别之间也显示出重叠。我们的结果表明,SAEs以分布式且依赖类别的方式局部化形态句法信息,而非通过原子语法特征。

英文摘要

Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.

发表机构

  • University of Pisa(比萨大学)

机构由 AI 辅助整理,请以论文原文为准。

↑