发表机构
Gradient PG; Gdańsk University of Technology(Gradient PG; 格但斯克理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过稀疏探针和激活修补,识别出冻结BERT中支持AI文本检测的稳定神经元子集,并发现指令微调模型在最后一层集中更多相关神经元,且该子集可跨生成器泛化。
AI 中文摘要
AI生成的文本检测器在标准基准上取得了高准确率,但驱动这些预测的内部表征仍鲜为人知。我们研究了冻结的BERT-base-uncased编码器中哪些神经元支持AI文本检测,使用了涵盖六个生成器的RAID基准,这些生成器跨越纯基础模型和指令微调模型。我们将Gurnee等人(2023)的L1-to-L2稀疏探针协议应用于全部9,216个CLS隐藏状态维度(12层×768),我们称之为神经元。该过程为每个生成器恢复了一个稳定的、少于1%的神经元集合,该集合在折叠和种子间保持一致;仅限于该集合的探针保留了大部分全特征检测准确率。双向激活修补证实了该集合的因果相关性:在两个方向上,它翻转预测的频率比大小匹配的随机集合高一个数量级。对相同神经元进行均值消融后,准确率基本保持不变;因此信号是冗余分布的。跨生成器分析揭示了一种二分结构:指令微调生成器将30-36%的稳定神经元集中在BERT的最后一层,而两个基础生成器均低于14%,这与后训练对齐的第12层足迹一致。留一簇族评估显示,所选神经元在未见生成器族上保留了全特征上限的86-94%,因此检测器可以在一个小的固定子空间上运行,而无需为每个生成器重新识别神经元。
英文摘要
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed. Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30-36% of stable neurons in BERT's final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment. Leave-one-family-out evaluation shows the selected neurons retain 86-94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.
CommentsAccepted to EMNLP 2026 (Main Conference)