通过非局部性测量SAE特征的语义抽象性
Measuring Semantic Abstractness of SAE Features via Nonlocality
浏览论文内容
中文总结 AI 辅助
该研究提出特征非局部性(FNL)指标,可区分SAE特征的抽象层级,用于审计越狱缓解特征及引导DeepSeek-R1模型特征提升数学推理准确率。
中文摘要 AI 辅助
稀疏自编码器(Sparse Autoencoders, SAEs)已通过理解与任务相关且具有因果效应的特征,为大型语言模型(LLM)的推理、越狱等行为提供了机制性解释。为评估这类机制性解释,后续研究必须区分表层词汇特征与真正的高层特征。然而,基于自动解释(autointerp)的语义描述或因果引导效用均无法完全解决特征的抽象层级问题。为此,我们提出特征非局部性(Feature Nonlocality, FNL),其定义为SAE特征激活的归一化逐位置影响的熵。研究表明,FNL与现有基于LLM的特征语义抽象性代理指标相关,且能成功区分依赖上下文的推理特征与基于token的特征:在由一个上下文特征和一个token级特征组成的随机配对中,73%至84%的情况会将更高的FNL分配给上下文特征。我们展示了两项下游应用:对用于缓解越狱的SAE特征进行审计,发现最有效的特征多为低FNL的位置特征,而非真正识别有害意图的特征;在DeepSeek-R1-Distill-Llama-8B中引导高FNL特征,使MATH-500准确率较未引导模型提升4.6个百分点,且优于引导低FNL特征的效果,尽管该提升具有模型特异性。我们得出结论,FNL为SAE特征的抽象层级提供了一种不依赖LLM、无标签、具有相关性的佐证,可用于评估机制性解释及为下游干预选择特征。
英文摘要
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in $73$--$84\%$ of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by $4.6$ points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.