arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24645cs.LGcs.AIcs.CL

稀疏自编码器对概念和函数进行编码:特征效应的下游几何结构

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty, Iryna Gurevych, Subhabrata Dutta

首次发表
浏览论文内容

中文总结 AI 辅助

研究稀疏自编码器特征与模型行为不一致问题,引入FEGA框架分析特征干预致模型对数几率变化的几何结构,区分值类和指针类特征,揭示特征可解释性及因果相关性与引导方向无关的现象。

中文摘要 AI 辅助

稀疏自编码器(SAE)作为可解释性工具的广泛应用受到SAE特征与模型行为之间不一致联系的限制。具有清晰激活描述的特征可能具有微弱或意外的因果效应;引导可能因提示而异或与预期方向相反;基于激活的特征选择可能会遗漏产生所需输出变化的特征。以往工作研究模型内部特征的几何结构,而本文研究特征干预导致的模型对数几率变化的几何结构。引入特征效应几何分析(FEGA),一个无监督框架,去除不同上下文中相同激活的SAE特征并分析对数几率变化云。结果显示跨SAE变体,一致的一维效应罕见。为解释这种变化,区分了与事实属性等静态信息相关的值类特征和与上下文相关操作相关指针类特征。值类特征更常表现出结构化、低维效应,指针类特征则主要表现出扩散效应。结果表明一个特征无需提供稳定引导方向也可具有可解释性和因果相关性。

英文摘要

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

↑