D分数:一种用于大语言模型中幻觉检测的谱隐藏状态信号
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
浏览论文内容
中文总结 AI 辅助
研究大语言模型幻觉检测,提出基于隐藏激活几何结构的D分数,通过单次前向传播计算,用作幻觉分数。经实验验证,该分数是强大的隐藏状态信号,检测时无需外部验证器等,为幻觉检测提供新方法。
中文摘要 AI 辅助
大语言模型可能生成错误、无证据支持或与模型内部表示不一致的流畅文本。我们从隐藏激活的几何结构研究幻觉检测,并引入D分数,它是通过单次前向传播计算出的简单谱统计量。对于固定模型、层和容差参数,D分数计算隐藏激活矩阵中奇异值接近主导奇异值的奇异方向数量。我们将此量用作幻觉分数,当输入文本的D分数大于预定义量时将其分类为幻觉。动机是当模型处理与自身内部状态信息冲突的文本时,隐藏表示可能编码断言内容和某种反证据等,导致隐藏轨迹分布在更多奇异方向。我们通过轻量级谱论证形式化此直觉,并在FAVA注释和RAGTruth上评估结果检测器。实验表明D分数是用于幻觉检测的强大隐藏状态信号,无需外部验证器、检索步骤和多轮生成。
英文摘要
Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that conflicts with information available in its own internal state, the hidden representation may encode both the asserted content and some form of counter-evidence, uncertainty, correction, or lack of support; this can make the hidden trajectory spread across additional singular directions. We formalize this intuition through a lightweight spectral argument and evaluate the resulting detector on FAVA-Annotation and RAGTruth. The experiments indicate that the D-Score is a strong hidden-state signal for hallucination detection, while requiring no external verifier, no retrieval step, and no multiple generations.
发表机构
- Department of Computer Science and Engineering, University of Bologna(博洛尼亚大学计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。