发表机构
Korea University; Konkuk University(高丽大学; 建国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现LLM幻觉检测的隐藏状态信号以单一均值漂移分量为主,简单的L2正则化逻辑回归等方法即可达到良好性能,LayerMix可聚合信号实现接近先知层的检测效果。
AI 中文摘要
隐藏状态探测方法可有效检测大语言模型(LLM)的幻觉问题,但该信号的几何特征仍未得到充分表征,这推动了探测架构日益复杂。在采用配对示例范式的三个7B规模模型和三个数据集上,研究发现该信号主要由单一均值漂移分量主导,移除该方向会使检测性能降至随机水平。收缩线性判别分析可缩小一维分类器与全维分类器之间约73%的差距,因此明显的架构复杂性主要反映了高维协方差估计的难度,而非存在可利用的非线性特征。简单的L2正则化逻辑回归(AUROC为0.952)的性能与十二种受控架构替代方法相当或更优,且在匹配范式下,研究的多层聚合方法优于CLAP跨层注意力探测方法。由于该信号跨越连续层带,LayerMix方法可聚合该信号,无需访问先知层即可达到先知层的性能。本研究在受控配对示例范式内对上述几何特征进行了表征,代码可在指定URL获取。
英文摘要
Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.
Comments19 pages, 7 figures, 20 tables. Accepted to EMNLP 2026 (Main Conference). Code: https://github.com/js-lee-AI/LayerMix