发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
我们提出并验证了语言模型表示中存在原子特征的理论,通过恢复原则预测并证实了SAE特征随规模增大的稳定性与共享性,支持了扩展SAE的可行性。
AI 中文摘要
我们发展并检验了一种关于语言模型表示的理论,该理论认为其中存在原子特征。我们的主要理论洞见是,在这样的模型中,规模递增的稀疏字典(例如SAE)会恢复训练数据中最普遍原子的递增前缀。这一“恢复原则”产生了三个可检验的预测:小型SAE中的许多特征被所有更大的SAE共享,在不同数据上训练的SAE共享两者中都普遍存在的特征,并且足够大的SAE能同时恢复父特征和子特征。与SAE特征不稳定且随规模增大而“分裂”的传统观点相反,我们发现这些预测在规模从512到131,072的SAE上成立,这些SAE在两个大型嵌入模型上训练。从理论角度看,我们的结果暗示了基于原子特征的表示科学理论的前景。从实践角度看,我们的结果暗示了扩展SAE的前景。
英文摘要
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.