超越词元尺度:用于可靠语义特征发现的块级稀疏自编码器
Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
- The University of Hong Kong(香港大学)
- Hunyuan Team, Tencent(腾讯混元团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出块级稀疏自编码器(Mean-Chunk、Cross-Chunk、Joint-Chunk),通过编码块级均值池化激活来发现可靠语义特征,在检索、推理检测和引导等下游任务中取得改进。
AI中文摘要:
稀疏自编码器(SAE)能够揭示有助于我们理解和引导语言模型的特征,但忠实的重建并不能保证产生有信息量的概念。词元级目标在鼓励语义内容的同时也奖励词汇和格式细节,所有这些都在争夺有限的稀疏预算。我们提出了一族块级SAE,它们对块(即连续的词元跨度)上的均值池化激活进行编码:Mean-Chunk重建观察到的块,Cross-Chunk预测一个独立处理的相邻块,而Joint-Chunk则结合这两个目标。这些设计将更大观察单元的影响与预测跨段落共享信息的影响分离开来。在匹配的训练数据下,块级SAE仍然是强大的可解释性工具,同时学习到能够捕获高层概念并对相关内容做出选择性响应的可靠语义特征。它们的优势是互补的:Mean-Chunk改进了高层特征发现、超越表面线索的推理检测以及引导;Cross-Chunk在文档检索和分类迁移方面领先,同时产生选择性的、持久的特征。改变SAE所看到和预测的内容,能够为更有意义的任务产生可靠的语义特征。我们通过在下游任务(如检索、推理检测和引导)中的收益展示了其实用价值。
英文摘要:
Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of chunk-level SAEs that encode mean-pooled activations over chunks, each a contiguous span of tokens: Mean-Chunk reconstructs the observed chunk, Cross-Chunk predicts an independently processed neighbor, and Joint-Chunk combines both targets. These designs separate the effect of a larger observation unit from that of predicting information shared across passages. With matched training data, chunk-level SAEs remain powerful interpretability tools while learning reliable semantic features that capture high-level concepts and respond selectively to relevant content. Their strengths are complementary: Mean-Chunk improves high-level feature discovery, reasoning detection beyond surface cues, and steering; Cross-Chunk leads document retrieval and classification transfer while producing selective, persistent features. Changing what an SAE sees and predicts yields reliable semantic features for more meaningful tasks. We demonstrate their practical value through gains across downstream tasks such as retrieval, reasoning detection, and steering.