可检测性差距:语言模型幻觉检测中的隐藏异质性
The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
浏览论文内容
中文总结 AI 辅助
本研究揭示语言模型幻觉检测中基于一致性的方法存在可检测性差距,不同模型和数据集上高、低一致性状态差异显著,表明聚合指标掩盖了模型依赖的异质性,需采用状态条件评估。
中文摘要 AI 辅助
基于采样的一致性方法被广泛用于幻觉检测,然而聚合性能可能掩盖哪些错误可被检测的系统性差异。本研究在四个语言模型和三个事实性问答数据集上考察了这种异质性。通过答案一致性对幻觉进行划分,揭示了高一致性(Ghost)和低一致性(Flickering)两种状态,并存在明显的可检测性差距,AUC相差0.35至0.46。由于用于定义状态和衡量该差距的统计量高度耦合(|ρ|≈0.94至1.00),原始结果被视为基于一致性的检测方法本身的性质,而非独立证据。在固定状态分配后,词汇和语义响应离散度保持了这种不对称性,所有12个模型和数据集设置中的bootstrap 95%置信区间均不包含零。一项更严格的测试使用单个扩散轨迹且不包含跨种子信息,在全部三个LLaDA数据集上保持了这种不对称性(p<0.005),并在全部三个Dream数据集上方向性保持一致,其中一个达到显著水平。困难状态在不同模型中的普遍性差异很大(16%至77%),且匹配的提示在不同模型之间经常改变状态。这些发现表明,聚合检测指标掩盖了语言模型失败中持续存在的、依赖于模型的异质性,并促使采用基于状态的评估方法。
英文摘要
Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreement reveals high agreement (Ghost) and low agreement (Flickering) regimes with an apparent detectability gap of $0.35$ to $0.46$ AUC. Because the statistics used to define the regimes and measure this gap are strongly coupled ($|ρ|\approx0.94$ to $1.00$), the raw result is treated as a property of agreement based detection rather than independent evidence. After freezing regime assignments, lexical and semantic response dispersion preserve the asymmetry, with bootstrap $95\%$ intervals excluding zero in all $12$ model and dataset settings. A stricter test using individual diffusion trajectories and no cross seed information preserves the asymmetry across all three LLaDA datasets ($p<0.005$) and directionally across all three Dream datasets, with one reaching significance. The hard regime varies substantially in prevalence across models ($16\%$ to $77\%$), and matched prompts frequently change regimes between models. These findings show that aggregate detection metrics conceal persistent, model dependent heterogeneity in language model failures and motivate regime conditioned evaluation.
发表机构
- R.V. College of Engineering(R.V. 工程学院)
- Technical University of Applied Sciences Würzburg-Schweinfurt(维尔茨堡-施韦因富特应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。