arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15804cs.CL

结合输入侧证据对齐的幻觉跨度检测

Hallucination Span Detection with Input-Side Evidence Alignment

Miyu Yamada, Yuki Arase

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型语言模型生成文本的幻觉问题,提出结合输入侧证据对齐的幻觉跨度检测任务,训练编码器模型通过预测置信度检测幻觉并对齐输入证据,实验验证了方法的有效性。

中文摘要 AI 辅助

幻觉(hallucination)仍是阻碍大型语言模型(LLMs)在条件文本生成中可靠应用的主要障碍。现有方法主要评估整个生成文本的事实性,对哪些输出跨度存在幻觉、或它们与输入的关联提供的信息有限。我们提出结合输入侧证据对齐的幻觉跨度检测任务,该任务可同时识别幻觉跨度并将输出标记与对应输入证据对齐。我们的方法基于以下观察:忠实的输出标记可由输入预测,而幻觉标记则无法预测。因此,我们训练一个基于编码器的模型,从输入表示中预测被掩码的输出标记,利用预测置信度进行幻觉检测,同时自然生成与输入的对齐。实验表明,所提方法能有效检测幻觉跨度并识别有意义的输入侧证据,人工评估也证实了预测对齐的质量。

英文摘要

Hallucinations remain a major obstacle to the reliable use of large language models (LLMs) in conditional text generation. Existing methods primarily assess the factuality of an entire generated text, providing limited insight into which output spans are hallucinated or how they relate to the input. We introduce the task of hallucination span detection with input-side evidence alignment, which jointly identifies hallucinated spans and aligns output tokens with the corresponding input evidence. Our approach is based on the observation that faithful output tokens are predictable from the input, whereas hallucinated tokens are not. We therefore train an encoder-based model to predict masked output tokens from the input representation, using prediction confidence for hallucination detection while naturally producing alignments to the input. Experiments show that the proposed method effectively detects hallucinated spans and identifies meaningful input-side evidence. Human evaluation confirms the quality of the predicted alignments.

发表机构

  • Institute of Science Tokyo(东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

↑