发表机构
University of Bern; Nanyang Technological University; Griffith University(伯尔尼大学; 南洋理工大学; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散语言模型的幻觉问题,现有方法有局限。本文提出无训练的TRE指标,通过在模型解码过程中沿时空维度提取熵信号来估计幻觉风险,经实验验证其性能有竞争力,具有泛化性、效率和鲁棒性等优点。
AI 中文摘要
扩散大语言模型(D-LLMs)近来备受关注,但其可靠性受到幻觉问题的显著阻碍。现有检测方法主要基于训练范式,依赖数据驱动训练优化检测器,存在局限性。本文提出TRE,一种无训练的幻觉检测指标,它通过单代熵信号直接估计幻觉风险,无需检测器训练或重复采样。TRE在D-LLM解码过程中沿时空维度提取熵信号,从令牌级空间和扩散步骤级时间角度获取信号并聚合得到TRE。实验表明TRE性能有竞争力,具备强泛化性、效率和鲁棒性。
英文摘要
Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination problem. Existing hallucination detection approaches for D-LLMs mainly follow a training-based paradigm, relying on data-driven training to optimize the detector. Such reliance not only limits their generalizability across domains models but also incurs additional training cost and deployment overhead. To address these limitations, we propose TRE, a training-free hallucination detection metric for D-LLMs. TRE is a parameter-free and single-run metric that estimates hallucination risk directly from the entropy signals of a single generation, without requiring any detector training or repeated sampling. TRE extracts entropy signals within the D-LLM decoding process along both the spatial and temporal dimensions. From a token-level spatial perspective, we focus on revealing tokens as the most informative carriers of uncertainty, capturing where uncertainty is actively committed. From a diffusion step-level temporal perspective, we empirically identify the dominance of late-step entropy and hence aggregate these signals with a simple linear weighting scheme to obtain TRE. Extensive experiments on multiple D-LLMs and QA datasets demonstrate that TRE achieves competitive performance, while enjoying strong generalizability, efficiency, and robustness.
Comments25 pages