幻觉神经元及其发现位置:关于幻觉神经元存在性的研究
Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons
查看机构详情
- Trakya University(特拉基亚大学)
- DRIVE, Great Ormond Street Hospital for Children NHS Foundation Trust(DRIVE,大奥蒙德街儿童医院NHS基金会信托)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出五步诊断协议,验证LLM中幻觉神经元的检测与定位,发现稀疏预测结构可与非唯一神经元选择共存,强调常规诊断验证的必要性。
中文摘要 AI 辅助
大型语言模型(LLM)的可解释机器学习日益依赖于稀疏探测方法,这些方法识别出据称能够检测并因果影响事实性回忆、安全对齐和幻觉等行为的小规模神经元集合。这些主张对模型审计和行为引导具有重要意义,然而它们很少针对$L_1$正则化探测在相关高维特征空间中的已知失效模式进行测试。我们提出一个五步诊断协议,涵盖特征相关性、自举稳定性、稀疏与密集排序不一致性、干预基线和跨数据集评估,作为稀疏神经元定位主张的最低标准。我们使用所提出的方法考察了先前的工作,特别是在TriviaQA、BioASQ和NQ-Open数据集上使用开源LLM对H神经元的研究。我们的结果表明,检测在两个模型和数据集上均可复现,并且超过了TriviaQA和BioASQ数据集上原始报告的AUROC差距。Gemma 3 4B在匹配数据集上始终优于MedGemma 4B,在TriviaQA上的AUROC差距为+0.311对+0.235,在BioASQ上为+0.474对+0.455,在NQ-Open上为+0.128对+0.112。在$n = 500$和五个随机种子下的因果验证显示,超出随机同层基线的统计显著效应。同时,诊断结果表明所选神经元并非唯一定位。在三个Gemma 3 4B设置中,22个选定的H神经元中有19个与其他特征的Pearson $|r| > 0.7$,自举选择仅显示中等稳定性,稀疏和密集排序仅弱重叠。我们的发现表明,稀疏预测结构可以与非唯一神经元选择共存。在机械可解释性中,常规诊断验证对于区分检测主张与定位主张是必要的。
英文摘要
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing and behavioral steering, yet they are rarely tested against known failure modes of $L_1$-regularized probing in correlated, high-dimensional feature spaces. We propose a five-step diagnostic protocol covering feature correlation, bootstrap stability, sparse versus dense ranking disagreement, intervention baselines, and cross-dataset evaluation as a minimum standard for sparse-neuron localization claims. We investigate prior work using our proposed approach, specifically on H-neurons using open-source LLMs across TriviaQA, BioASQ, and NQ-Open datasets. Our results demonstrate detection replicates across both models and datasets, and exceeds the original reported AUROC gaps for TriviaQA and BioASQ datasets. Gemma 3 4B consistently outperforms MedGemma 4B on matched datasets, with AUROC gaps of +0.311 versus +0.235 on TriviaQA, +0.474 versus +0.455 on BioASQ, and +0.128 versus +0.112 on NQ-Open respectively. Causal validation at $n = 500$ with five random seeds shows statistically significant effects beyond random same-layer baselines. At the same time, the diagnostic results indicate that the selected neurons are not uniquely localized. Across the three Gemma 3 4B settings, 19 of 22 selected H-Neurons have Pearson $|r| > 0.7$ with other features, bootstrap selections show only moderate stability, and sparse and dense rankings overlap only weakly. Our findings show that sparse predictive structure can coexist with non-unique neuron selection. Routine diagnostic validation is necessary to distinguish detection claims from localization claims in mechanistic interpretability.