发表机构
Rutgers University; NVIDIA Research(罗格斯大学; 英伟达研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出参考特征图谱用于语言模型审计,通过拟合线性解码器实现,有图谱和残差两个互补通道。在多模型上训练并审计米斯特拉尔和文心一言目标,残差通道能控制植入机制,还揭示文心一言相关集群,为模型审计提供新方法。
AI 中文摘要
审计新的语言模型通常意味着从头重新学习和重新解释其内部特征。我们提出了一种参考特征图谱:一个在参考面板上训练一次并用于新目标的稀疏特征库,通过仅拟合线性解码器来附加。这产生了两个互补视图。图谱通道在已解释的面板特征上读取目标,提供跨模型的稳定坐标系。残差通道仅从图谱未能重建的部分学习特征,使“参考面板之外”成为明确的审计信号。我们在五个7 - 9B指令调整模型上训练留一法图谱,并审计留出的米斯特拉尔和文心一言目标。在注入两个目标的三个受控LoRA隐藏目标上,残差通道使植入机制在运行时完全可控,而匹配的对照不受影响,并在两个目标上都将植入目标恢复为排名最高的潜在目标;在米斯特拉尔上,针对逐对交叉编码器基线进行从头对头基准测试而重新训练每个目标的SAE和成对交叉编码器基线时,两者均未做到这一点。在文心一言- 2.5上,同一通道还揭示了一个与面板相关的政治框架集群;对其进行引导会改变审计的框架指标,而域外对照保持不变。
英文摘要
Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged.