发表机构
University of Technology Sydney; City University of Macau; Meta; actAVA AI; Microsoft AI; Squirrel Ai Learning; East China Normal University(悉尼科技大学; 澳门城市大学; Meta; actAVA人工智能公司; 微软人工智能; 松鼠人工智能学习公司; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HalluTracer是一种在模型生成答案前聚合各层真值证据的幻觉检测框架,在六个开源语言模型和五个基准上优于白盒基线,将幻觉检测转化为深度聚合问题。
AI 中文摘要
即使经过充分对齐的大语言模型也会自信地生成事实上不正确的文本,使得幻觉成为高风险部署中持续存在的可靠性风险。不过这些模型的内部表征中包含线性可分的真值信号。然而,现有的白盒检测器会将这些证据坍缩为孤立的组件或单一深度,丢弃了全前向传播过程中分布的判别信息。我们提出HalluTracer,这是一种检测框架,在模型生成任何答案token之前,读取并聚合全前向传播各层的真值证据。几何分析表明,各层信号弱相关,因此简单的深度平均可抑制层特定噪声并捕获几乎所有线性可访问信息。在六个开源语言模型和五个幻觉基准上,HalluTracer始终优于匹配的白盒基线,提升幅度为1到14个点。总体而言,我们的工作将幻觉检测从层选择问题重新定义为由真值信号的几何稀疏性支配的深度聚合问题。
英文摘要
Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they read and the rules used to combine them, the source of improved prediction becomes difficult to identify. We separate these choices and find that retaining more layers improves prediction even under fixed equal weighting. An exact Fisher-ratio decomposition explains why the additional benefit of linear reweighting is limited on these probe scores: information inlayer-wise differences largely overlaps with that captured by the depth mean. Estimating additional weights can then offset this small benefit when the data used to fit them are limited. These findings motivate our proposed method HalluTracer, which averages layer-wise probe logits to predict truthfulness before decoding. Across six models and four benchmarks, including TruthfulQA, HalluTracer achieves the highest area under the receiver operating characteristic curve (AUROC) in 23 of 24 model--benchmark pairs among the compared methods. The results support using evidence from across the network without requiring a correspondingly more flexible aggregation rule, clarifying the distinct roles of layer selection and weighting in pre-decoding detection.