发表机构
Algoverse AI Research(Algoverse AI 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出利用SAE解码器空间的几何属性(邻居密度和最大余弦相似度)在干预前预测特征转向成本,实验表明该信号在多种模型和SAE配置下有效,为特征可控性筛选提供了初步依据。
AI 中文摘要
使用 SAE 特征进行转向需要对每个特征进行系数调优,目前这需要干预扫描。我们提出一个问题:SAE 自身的属性(在任何前向传播之前即可计算)能否预测哪些特征转向成本低或高。我们表明,SAE 特征可转向性的变化部分可由解码器空间几何结构预测:邻居密度和与附近解码器方向的最大余弦相似度,这两者均可在任何干预之前从 SAE 权重矩阵计算得出,它们根据实现固定行为效果所需的转向量对特征进行排序(ρ 最高达 -0.546,p < 10^-6,AUROC 在 0.610-0.822 之间,跨条件;该信号基于排名,与网格离散性一致)。这种几何-可转向性关系在两种 Gemma-2 模型规模(2B 和 9B)、两种 SAE 宽度(16K 和 65K)中均得到复现,并且在 Llama-3.1-8B-Instruct 上可跨架构检测到(ρ = -0.266,n = 300)。在采用 BatchTopK SAE 的 Qwen3-8B 上,几何结构能预测特征是否完全可转向,但无法预测响应特征之间的连续排序,这揭示了与 SAE 训练机制相关的边界条件。在两种模型的深层比例层深度处,信号减弱,此时转向成本超过我们的干预预算,这构成了一致的深度边界。这些结果提供了初步证据,表明转向前的几何结构可以部分指导系数选择,为在部署前筛选特征的可控性提供了一条途径。
英文摘要
Steering with SAE features requires per-feature coefficient tuning, which currently demands intervention sweeps. We ask whether properties of the SAE itself, computable before any forward pass, predict which features will be cheap or expensive to steer. We show that variation in SAE feature steerability is partially predicted by decoder-space geometry: neighbor density and maximum cosine similarity to nearby decoder directions, both computable from the SAE weight matrix before any intervention, rank features by how much steering they require for a fixed behavioral effect ($ρ$ up to $-0.546$, $p < 10^{-6}$, AUROC 0.610-0.822 across conditions; the signal is rank-based, consistent with grid discreteness). This geometry-steerability relationship replicates across two Gemma-2 model scales (2B and 9B), two SAE widths (16K and 65K), and is detectable cross-architecturally on Llama-3.1-8B-Instruct ($ρ= -0.266$, $n = 300$). On Qwen3-8B with BatchTopK SAEs, geometry predicts whether a feature is steerable at all but not the continuous ordering among responsive features, revealing a boundary condition tied to SAE training regime. The signal weakens at deep proportional layer depth in both models, where the cost of steering exceeds our intervention budget, a consistent depth boundary. These results provide preliminary evidence that pre-steering geometry can partially inform coefficient selection, offering a path toward screening features for controllability before deployment.
Comments12 pages, 3 figures, 5 tables. Accepted at the Mechanistic Interpretability Workshop at ICML 2026 and AIW 2026 at COLM