arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

寻找语言模型中负责网络信息检索的注意力头与神经元

Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models

Md Abdul Kadir, Md Mohasin Hossain, Daniel Sonntag

arXiv 2610.08200首次发表:更新:

发表机构

University of Oldenburg; German Research Center for Artificial Intelligence (DFKI); Saarland University(奥尔登堡大学; 德国人工智能研究中心(DFKI); 萨尔兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果消融等方法,验证了语言模型中特定注意力头及部分神经元对网络信息检索的责任,发现跨模型有效,但责任是否集中于单个神经元因模型而异。

AI 中文摘要

我们探究是否特定的注意力头,以及更精细地,这些注意力头内部的特定神经元,负责识别语言模型的上下文是否包含网络基础设施信息(主机名及其IP地址配对),并且这种责任能否通过因果方式而非仅凭相关性得到验证。在注意力头层面,答案是肯定的,这跨越了五个模型、三个架构家族:在每个模型中,通过因果消融筛选发现并由匹配的负样本和无上下文对照进行选择性测试的一小组注意力头(在128到1152个候选中占1到9个),支持一个检测器,其留出准确率达到99.5%至100%。我们接着探究一个注意力头的责任是集中于单个神经元还是分散在其各个维度上;这因模型而异。在一个模型中,最顶部的注意力头的信号集中于单个神经元,这一发现独立地通过因果干预和相关性排序得到,两者完全一致(AUC=1.000,与完整注意力头匹配)。在另一个模型中,单个干净的注意力头作为一个整体工作(AUC=1.000),但其内部因果排名最佳的神经元却并非如此(AUC=0.665),因此那里的责任分散在整个注意力头中。其余三个模型则介于两者之间。在一个由不同机构收集的独立数据集上(反向DNS记录而非发现数据),每个模型的完整注意力头检测器标记了100%的正样本;单神经元版本迁移可靠性较低,在一个模型中得分低于随机水平。针对特定网络信息实体的因果注意力头查找在跨模型和架构时有效;这一发现能下探到单个神经元的程度各不相同,需要对每个模型进行检验。

英文摘要

We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100\% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100\% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.

Comments13 pages, 2 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑