MechaTerp-TRACE:语言模型组件消融分析的新方法
MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models
- National Institutes of Health, National Library of Medicine(美国国立卫生研究院、国家医学图书馆)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MechaTerp-TRACE方法,通过逐组件消融统一比较语言模型各组件对命名实体生成的因果贡献,发现实体知识局部化主要源于通用生成机制,影响知识编辑方法。
AI中文摘要:
大型语言模型的可解释性研究已产生了关于前馈层中事实回忆和自注意力中标记关系的解释,但很少有工作提供一种统一的方法来比较不同架构组件对模型输出的因果贡献。我们引入MechaTerp(机制可解释性套件)-TRACE(教师强制消融组件效应注册子集),这是一种架构和研究方法,用于衡量语言模型的每个注册组件在多大程度上支持命名实体的生成。TRACE每次消融一个组件,并测量在固定答案标记处输出分布的变化,从而可以在共同尺度上比较从整个Transformer块到单个神经元和输出logits的组件类型。我们将其应用于十三个指令微调的稠密解码器模型,涵盖五个模型家族和十亿至三百亿参数,在48个医学和42个通用知识提示中消融了49,656个组件。我们发现,在每一个模型中,携带最大效应的组件是相同的少数几个位置固定的组件,无论提示询问哪个实体,而且一旦这些组件被移除,在十三个模型中的十一个中,剩余的支持接近均匀分布。因此,实体知识的明显局部化在很大程度上归因于通用生成机制,这对假设实体知识位于可查找位置的方法(包括定向知识编辑)具有直接影响。
英文摘要:
Interpretability research on large language models has produced accounts of factual recall in feed-forward layers and of token relationships in self-attention, but little work offers a unified way to compare the causal contribution of different architecture components to a model's output. We introduce MechaTerp (the Mechanistic Interpretability suite) -TRACE (subset for Teacher-forced Registry of Ablated Component Effects), an architecture and study that measures how much each registered component of a language model supports the production of a named entity. TRACE ablates one component at a time and measures the resulting change in the output distribution at a fixed answer token, so component types from whole transformer blocks down to individual neurons and output logits can be compared on a common scale. We apply it to thirteen instruction-tuned dense decoder models spanning five families and one to thirty billion parameters, ablating 49,656 components across 48 medical and 42 general-knowledge prompts. We find that the components carrying the most effect are the same few, positionally fixed components in every model, regardless of which entity a prompt asks about, and that once these are removed, the remaining support is close to evenly spread in eleven of the thirteen models. Apparent localisation of entity knowledge is therefore largely attributable to generic generation machinery, which has direct consequences for methods that assume entity knowledge sits in a findable place, including targeted knowledge editing.