发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究神经网络激活可解释性方法,提出HARP方法,即让语言模型智能体利用激活向量数据库及工具,通过迭代查询、形成并验证假设来实现无需训练的可解释性,该方法在多方面表现出色,还凸显现有训练方法局限并推动新基准测试。
AI 中文摘要
神经网络激活的可解释性方法涵盖了广泛的成本范围,从低成本、无需训练的技术(如线性探针、主成分分析、奇异值分解)到更昂贵的基于训练的方法(如自联想编码器和激活预言机)。基于训练的方法通常更强大,部分原因是它们在训练期间利用了大型激活数据集。这就引出了一个自然的问题——它们是否真的能揭示超出训练数据集本身可恢复的见解?为了解决这个问题,我们为一个语言模型智能体配备了一个激活向量数据库及其文本上下文,以及用于操作激活的工具——在潜在空间中投影出方向、计算激活差异和平均值。该智能体迭代查询数据库,根据检索到的样本形成假设,并通过构建线性探针来验证它们。我们将这种方法称为HARP,即假设驱动的智能体检索与探测。尽管不涉及任何训练,但HARP在概念发现、概念检测、模型引导和秘密提取方面优于激活预言机和基于自联想编码器的智能体。无需训练的设计也使HARP成本更低、更灵活:只要现有数据集不足,就可以按需索引新数据集。更广泛地说,我们的结果表明,当前基于训练的方法尚未提取超出其训练数据的见解,并推动了明确要求可解释性方法展示此类见解的基准测试。我们在这个https URL上发布了我们的代码
英文摘要
Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training-based ones (such as SAEs and activation oracles). Training-based methods are typically more powerful, in part because they leverage large activation datasets during training. This raises a natural question - do they actually surface insights that go beyond what is recoverable from the training dataset itself? To address this, we equip an LLM agent with a vector database of activations paired with their textual contexts, along with tools for manipulating activations - projecting out directions in latent space, computing activation differences and averages. The agent iteratively queries the database, forms hypotheses from the retrieved samples, and validates them by constructing linear probes. We call this method HARP, for Hypothesis-driven Agentic Retrieval and Probing. Despite not involving any training, HARP outperforms both activation oracles and SAE-based agents on concept discovery, concept detection, model steering, and secret elicitation. The training-free design also makes HARP substantially cheaper and more flexible: new datasets can be indexed on demand whenever existing ones prove insufficient. More broadly, our results suggest that current training-based methods do not yet extract insights beyond their training data, and motivate benchmarks that explicitly require interpretability methods to demonstrate such insights. We release our code at https://github.com/SriramB-98/HARP
Comments25 pages, 7 figures