arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检测与修复检索增强生成中的幻觉现象

Detecting and Repairing Hallucinations in Retrieval-Augmented Generation

Sai Krishna Reddy Mulakkayala, Niki van Stein, Aske Plaat

arXiv 2608.29307首次发表:更新:

发表机构

Leiden University(莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对检索增强生成(RAG)的幻觉问题,借助RAGTruth基准数据集,对比三种修复策略的效果,发现策略在依据保留权衡上各有侧重,需依据答案有用性选择。

AI 中文摘要

语言模型越来越多地通过查阅检索到的文档而非仅依赖记忆来回答问题,这种设计如今在搜索助手和企业知识工具中已十分常见。将模型基于检索文本进行推理可减少无依据陈述,但无法彻底消除此类问题,且读者无法区分有依据的句子与编造的句子。多数针对该问题的研究止步于检测,然而标记错误答案对阅读者毫无帮助,且人们对标记后应采取何种后续行动知之甚少。本研究使用人工标注了无依据段落的基准数据集RAGTruth,将每个被标记的答案拆分为单个事实主张,对照检索来源逐一核查,并对比三种丰富度递增的修复策略与保留原答案的效果:删除无依据主张、用来源文本替换该主张、重写该主张。来自不同家族的三种语言模型对916个修复后的答案进行评判。结果显示,每种策略均能降低被评判为包含无依据内容的答案比例,且三位评判者对策略排序达成一致:删除策略降幅最大,但保留的原答案文本占比最低(64.3%);重写策略保留的文本占比最高(80.1%),降幅最小。修复操作不仅针对错误答案:83.5%被标注为干净的答案也被编辑过。这些策略在依据保留权衡上占据不同位置,而非形成质量排名,选择何种策略需有关答案有用性的证据,而自动指标无法提供此类证据。

英文摘要

Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants and enterprise knowledge tools. Grounding a model in retrieved text reduces unsupported statements but does not eliminate them, and a reader cannot tell a grounded sentence from an invented one. Most research on this problem stops at detection, yet flagging a faulty answer changes nothing for the person reading it, and little is known about which action should follow. Using RAGTruth, a benchmark whose unsupported passages are annotated by hand, we split each flagged answer into individual factual claims, check each against the retrieved source, and compare leaving the answer untouched with three repair strategies of increasing richness: deleting an unsupported claim, replacing it with source text, and rewriting it. Three language models from different families judge the 916 repaired answers. Every strategy reduces the proportion of answers judged to contain unsupported content, and all three judges agree on the ordering. Deletion achieves the largest reduction while retaining least of the original answer, at 64.3% of the text, whereas rewriting retains 80.1% and reduces least. Repair is not confined to faulty answers: 83.5% of answers annotated clean are edited too. The strategies occupy different points on a grounding preservation trade-off rather than forming a quality ranking, and choosing between them needs evidence about answer usefulness that automatic metrics cannot supply.

Comments15 pages, 3 figures, 7 tables. Submitted to BNAIC/BeNeLearn 2026 as a Type A paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑