提示嵌入探针(PEP):基于隐藏状态的大语言模型幻觉检测
Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
浏览论文内容
中文总结 AI 辅助
本研究提出提示嵌入探针(PEP),一种白盒幻觉检测方法,通过少量可学习提示嵌入扩充标准线性探针,在TriviaQA等数据集的Qwen3模型上验证其在生成前、跨模型等场景的有效性,仅需少量参数即可提升检测性能。
中文摘要 AI 辅助
大语言模型(LLMs)可生成流畅且有用的响应,但仍易产生幻觉。我们提出提示嵌入探针(Prompt Embedding Probes,PEP),一种从冻结LLM的隐藏状态中进行答案级幻觉检测的白盒方法。PEP通过用少量可学习的提示嵌入扩充输入,扩展了标准线性探针。我们在TriviaQA、GSM8K和MedQA数据集上,使用多个规模的Qwen3模型对PEP进行评估。在主要的分布内设置中,PEP比标准线性探针提升了基于隐藏状态的检测效果。我们进一步评估了PEP在生成前预测、跨模型迁移和分布外泛化中的表现:PEP在生成前和跨模型设置中仍保持有效,而稳健的跨数据集迁移仍存在困难。这些结果表明,基于提示的适配可在保持主干冻结且仅添加少量可训练参数的情况下,增强隐藏状态探针的性能。
英文摘要
Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP), a white-box method for answer-level hallucination detection from the hidden states of a frozen LLM. PEP extends standard linear probes by augmenting the input with a small number of learnable prompt embeddings. We evaluate PEP on TriviaQA, GSM8K, and MedQA using Qwen3 models at multiple scales. PEP improves hidden-state-based detection over standard linear probes in the main in-distribution setting. We further evaluate PEP for pre-generation prediction, cross-model transfer, and out-of-distribution generalization. PEP remains effective in the pre-generation and cross-model settings, whereas robust cross-dataset transfer remains difficult. These results show that prompt-based adaptation can strengthen hidden-state probing while keeping the backbone frozen and adding only a small number of trainable parameters.
发表机构
- HSE University(高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。