发表机构
University of Cambridge; University of Oxford(剑桥大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在语言模型中学习特征时间尺度的问题,提出持久稀疏自动编码器,通过为特征学习持久性系数扩展标准SAEs,实验表明其能保持竞争力的重建质量,为解释和监测语言模型带来新机会。
AI 中文摘要
稀疏自动编码器(SAEs)将语言模型激活分解为稀疏特征,但标准SAEs独立编码每个令牌,不暴露跨序列持续存在的信息。我们引入了持久稀疏自动编码器(Persistent SAEs),它通过为每个特征学习一个持久性系数来扩展标准SAEs,使模型能够学习哪些特征应该持续以及持续多长时间。我们的实验表明,它们在学习一系列特征时间尺度时保持了有竞争力的重建质量:快速特征表现为局部可解释的检测器,而慢速特征在持久状态下集中主题级信息。此外,如在提示注入监测案例研究中所示,慢速特征保留检测信号并在长上下文中保持因果有效性。这些结果表明,持久稀疏自动编码器通过持久语义表示为解释和监测语言模型开辟了新机会。
英文摘要
Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (slow) features accumulate information that identifies the current context. Moreover, we show in a prompt-injection monitoring case study that slow features preserve injection-related signals and remain causally effective over long contexts. These results suggest that Persistent SAEs offer new opportunities for interpreting and monitoring language models via persistent sparse features.