发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究SAE特征作为安全控制手柄的情况,引入匹配相干门控评估协议并应用于Gemma-2-9B-it,发现SAE特征消融有用范围窄,不同SAE top表现各异,结果表明应将基于SAE的安全干预作为依赖范围的控制机制评估。
AI 中文摘要
我们评估稀疏自动编码器(SAE)特征何时作为与安全相关行为的局部控制手柄。由于弱干预、不匹配的基线、模型鲁棒性或退化输出等原因,这个问题很难解决。我们引入了一种用于运行时安全干预的匹配相干门控评估协议,在匹配的目标效应点比较方法,主要目标指标仅在输出既被判定不安全又相干时计算有害合规性。将该协议应用于Gemma-2-9B-it的三个提示分割上,发现SAE特征消融的有用范围很窄。SAE top800以较低的总扰动和有竞争力的效用达到中低目标效应,但SAE top1600相对于匹配的密集拒绝方向基线失去效用,SAE top3200主要导致相干崩溃。人类审计证实相干门控消除了仅不安全的伪像,特征诊断表明有用范围由稳定的拒绝对齐特征头部驱动,其激活分离随秩迅速衰减。这些结果表明,基于SAE的安全干预应作为依赖于范围的控制机制进行评估,而不是假定为统一局部化。
英文摘要
We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.
Comments11 pages, 5 figures, 4 tables. Preliminary version; extended multi-model study in progress