arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27165cs.CL

预测的预测(PoP):用于大语言模型单次通过幻觉检测的层间激活融合

Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models

Himal Badu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出PoP机制,通过融合单次前向传播的层间隐藏表示检测大语言模型幻觉,在TruthfulQA基准上AUROC达75.5%,延迟增加不足1.2%,无需额外生成传播。

中文摘要 AI 辅助

自回归大语言模型(LLMs)常会以高解码置信度生成事实上错误的输出,这限制了它们在高风险工作流中的应用。现有的输出阶段不确定性指标在模型对错误断言过度自信时可能失效,而多样本验证流程则会引入大量内存与延迟开销。本研究探究生成过程中的内部隐藏状态转换动态是否能在无需辅助解码调用的情况下指示事实错误。我们提出预测的预测(Prediction of Prediction,PoP)机制,该机制通过在单次前向传播中融合不同深度的中间隐藏表示来捕获层转换不确定性。在TruthfulQA基准上使用自回归Transformer主干进行评估时,PoP在事实正确性分类任务中达到了受试者工作特征曲线下面积(AUROC)为75.5%的性能。该机制在基础前向传播内运行,增加的运行时延迟不足1.2%,且无需额外的生成传播。数值结果来自作者验证的实验实现,且受下文所述评估范围的限制。

英文摘要

Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows. Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions, while multi-sample verification pipelines introduce substantial memory and latency overhead. This work evaluates whether internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls. We introduce Prediction of Prediction (PoP), a mechanism that captures layer-transition uncertainty by fusing intermediate hidden representations across depth during a single forward pass. Evaluated on the TruthfulQA benchmark using autoregressive transformer backbones, PoP achieves an area under the receiver operating characteristic curve (AUROC) of 75.5% for factual-correctness classification. The mechanism operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes. The numerical results are reported from the author-verified experimental implementation and are bounded by the evaluation scope described below.

补充信息

↑