SharedSAE:跨语言模型的单一特征字典
SharedSAE: One Feature Dictionary Across Language Models
浏览论文内容
中文总结 AI 辅助
SharedSAE可跨语言模型共享特征字典,仅归一化选择分数并采用模型dropout实现单模型推理,在1B规模模型上训练后,保留96.6%的专用SAE解释方差,新模型可高效适配并复用潜在描述。
中文摘要 AI 辅助
稀疏自编码器(SAEs)被广泛用于解释语言模型的激活,但SAE的训练和潜在标签标注通常需针对每个模型重复进行。本文提出的SharedSAE可替代一组专用的单模型SAE,其结合了共享字典与模型特定的编码器-解码器对。与最近的方法不同,该方法仅归一化选择分数以保留激活幅值,并采用模型 dropout 实现单模型推理,无需在推理时使用所有模型。研究人员在四个涵盖不同家族和分词器的1B规模基础语言模型上训练SharedSAE,尽管其潜在表示跨模型共享,仍保留了专用SAE 96.6%的平均解释方差;其潜在激活表现出的跨模型相关性,是事后对齐的独立SAE的1.8倍,且潜在描述可跨模型迁移。字典冻结后,新模型可高效适配该字典,在复用共享潜在描述的同时,达到接近专用SAE的重建质量。
英文摘要
Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of dedicated per-model SAEs. Our method, SharedSAE, combines a shared dictionary with model-specific encoder-decoder pairs. Unlike the closest prior method, which discards activation magnitudes and requires all models at inference, SharedSAE instead normalizes only selection scores, preserving magnitudes, and uses model dropout for single-model inference. We train SharedSAE on four 1B-scale base language models spanning distinct families and tokenizers. Despite sharing its latents across models, SharedSAE retains 96.6% of dedicated SAEs' mean explained variance; its latent activations exhibit cross-model correlations 1.8 times as high as separate SAEs aligned post-hoc, and its latent descriptions transfer across models. After the dictionary is frozen, new models can be efficiently adapted to it, achieving near-dedicated-SAE reconstruction quality while reusing the shared latent descriptions.
发表机构
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
- Sorbonne Université(索邦大学)
- Tohoku University(东北大学)
- RIKEN AIP(理化学研究所人工智能项目)
机构由 AI 辅助整理,请以论文原文为准。