arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检索增强生成系统为何会产生幻觉:基于知识间隔金丝雀的惩罚感知评估

Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

Alden Do Rosario, Hussein Younes, Felipe Pires

arXiv 2608.26385首次发表:更新:

AI 中文总结

该研究针对RAG系统的幻觉问题,提出惩罚感知评估框架,结合不对称评分、知识间隔金丝雀和失败归因流程,在SimpleQA-Verified数据集上揭示系统差异源于不当作答,而非准确率,相关资源已公开。

AI 中文摘要

基于数量的准确率会奖励检索增强生成(RAG)系统进行猜测:一个对所有问题都作答的系统,会比知识库无法支撑答案时选择弃权(不执行)的系统表现更好。基于Kalai等人(2025)的置信度目标分析,我们提出了一种针对已部署RAG产品的惩罚感知评估框架,该框架结合了三部分内容:(i)不对称评分(回答正确得+1分,回答错误扣-4分,弃权(不执行)得0分);(ii)知识间隔金丝雀,即答案可被验证为不存在于知识库中的问题,因此任何对这类问题的回答都构成了来自参数记忆的无依据生成;(iii)失败归因流程,可区分检索、生成和弃权(不执行)策略的失败。将该框架应用于三个商业RAG系统及一个无检索基线模型,在SimpleQA-Verified数据集(1000个问题×3次重复,由跨领域三人评审团盲评,一致度达98.9%)上,我们发现各系统作答时的准确率较为接近(97.0%-98.0%),而金丝雀违规率则相差约6倍(16.7%对98.1%)。各系统的差异更多体现在它们是否会在不该作答时作答,而非正确回答的内容,惩罚感知评分也相应地重新排序了基于数量的排名;该重新排序在惩罚设置k=1到k=9时保持稳定。所有代码、配置、转录本及评审投票均已发布,供独立审计。

英文摘要

Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kalai et al. (2025), we present a penalty-aware evaluation framework for deployed RAG products, combining (i) asymmetric scoring (correct +1, wrong -4, abstain 0), (ii) knowledge-gap canaries, questions whose answers are verifiably absent from the knowledge base, so that any answer constitutes ungrounded generation from parametric memory, and (iii) a failure-attribution pipeline that separates retrieval, generation, and abstention-policy failures. Applying the framework to three commercial RAG systems and a no-retrieval baseline on SimpleQA-Verified (1,000 questions x 3 repeats, graded blind by a cross-family three-judge panel with 98.9% unanimity), we find that accuracy when answering is closely clustered across systems (97.0-98.0%), while canary violation rates differ roughly sixfold (16.7% vs. 98.1%). The systems are separated less by what they answer correctly than by whether they answer at all when they should not, and penalty-aware scoring reorders the volume-based ranking accordingly; the reordering is stable across penalty settings from k=1 to k=9. All code, configurations, transcripts, and judge votes are released for independent audit.

Comments13 pages, 1 figure, 5 tables. Code, full per-request logs, and all judge votes: https://github.com/adorosario/why-rags-hallucinate

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑