arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeuroActiSep:单次前向传播中从前馈神经元检测事实性幻觉

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

Ali Derogar Odolou, Reza Nazari, Mostafa Salehi

arXiv 2609.14448首次发表:更新:

发表机构

University of Tehran(德黑兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出NeuroActiSep方法,通过自定义数据集排序前馈神经元,迁移至其他问答数据集训练幻觉分类器,证明其性能与内部状态探针相当,并分析神经元分布与层深度影响。

AI 中文摘要

大语言模型中的幻觉降低了其可靠性并减缓了其采用速度。各种白盒研究已利用内部表示来检测真实性和事实性的模式。一种较少被研究的方法是识别与幻觉相关的前馈神经元。我们提出了一种方法,使用自定义的神经元选择数据集,对最终提示词标记处的前馈神经元进行排序。我们将选定的神经元身份迁移到其他事实性问答数据集上训练幻觉分类器。我们的工作提供了经验证据,表明使用选定神经元特征训练的探针与基于内部状态训练的探针性能相当。我们还分析了选定神经元的分布以及层深度对检测性能的影响。

英文摘要

Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated with hallucination. We propose a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset. We transfer the selected neuron identities to train hallucination classifiers on other factual question answering datasets. Our work provides empirical evidence that probes trained using the features from the selected neurons perform on par with probes trained on internal states. We also analyze the distribution of selected neurons and the effect of layer depth on detection performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑