发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SteerCheck是一项预先注册的归因审计方法,用于检测激活引导的对齐泄漏,揭示了Qwen3-14B等模型在激活引导中的特异性局限,为审计相关结论提供了可操作的框架。
AI 中文摘要
激活引导可在不确认其效果对预期概念具有特异性的情况下改变模型行为。本文提出SteerCheck,这是一项预先注册的归因审计方法,用于匹配非目标KL散度,并分离均值、受保护尾部、极性、迁移及语义主张。对960个Qwen3-14B干预的精确复现显示,常见控制方法存在互补性局限:各向同性方向占据狭窄的近正交区域,而符号随机化的同结构方向通常保留大量目标对齐。效果与符号随机化族内的符号余弦值强相关(ρ=.94);其25.3%的抽样值超过余弦值0.5,所有超过观测均值效果的抽样值余弦值均高于0.80。这种对齐泄漏本身并不使条件随机化检验失效,但限制了比较器可区分的内容,并促使研究者报告可交换性假设、构造诊断A及经验余弦分布。Qwen的完整门限保持为负,因受保护尾部在所有族中均不通过。在独立数据上,仅Qwen中存在连续边际迁移,无选定单元存在准确率迁移。预先注册的语言控制方法在Qwen和DeepSeek中通过完整门限,而通过的DeepSeek detox比较器排除了类别分离;所有名义上的通过对Γ=1.10敏感。冻结的三评估者开放生成评估支持DeepSeek中的事实校正,但不支持Qwen;自动评估器的校准失败(宏F1值为.562),因此空范围语义结果仍为描述性。SteerCheck使这些条件及混合结论具备可审计性。
英文摘要
Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomized family ($ρ=.94$); $25.3\%$ of its draws exceed cosine $.5$, and every draw exceeding the observed mean effect has cosine above $.80$. This alignment leakage does not by itself invalidate a conditional randomization test; it limits what the comparator can distinguish and motivates reporting exchangeability assumptions, a construction diagnostic $A$, and the empirical cosine distribution. The primary Qwen complete gate remains negative because the protected tail fails all families. On independent data, continuous margin transfers only in Qwen and accuracy transfers in no selected cell. Prospectively registered language controls pass the complete gate in Qwen and DeepSeek, while a passing DeepSeek detox comparator rules out categorical separation; all nominal passes are sensitive to $Γ=1.10$. Frozen three-rater open-generation evaluation supports factual correction in DeepSeek but not Qwen; the automatic judge fails calibration (macro-F1 $.562$), so null-wide semantic results remain descriptive. SteerCheck makes these conditional and mixed conclusions auditable.