arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

欺骗性基础:临床检索增强生成中的实体归因失败

Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation

Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim

arXiv 2607.09349首次发表:更新:

发表机构

Lunit(鲁尼特)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究临床检索增强生成中的欺骗性基础问题,通过对13个模型测试发现DG率8%-87%,医学等微调模型高达86.7%。控制消融确定机制,实体归因验证可检测DG,现有框架未实施,为该领域研究提供新视角和方法。

AI 中文摘要

检索增强生成评估检查模型声明是否在检索到的文档中有事实依据,但不检查检索到的证据是否归因于正确的实体。临床RAG响应可能通过所有自动检查,但将药物Y的临床证据呈现为关于查询药物X的证据,即欺骗性基础(DG)。通过对13个模型的控制因子基准测试,发现对抗条件下DG率在8%-87%之间。医学和生物医学微调模型高达86.7%,领域专业化加剧了失败。控制消融确定了机制,去除检索文档中特定实体的临床证据可消除实体归因失败。在部署的RAG系统中,740个药物-疾病对的生产测量发现总体DG为7.8%,新批准药物升至13.6%。实体归因验证检测DG的精度为97.0%,召回率为98.7%,现有框架未实施。

英文摘要

Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated check (zero hallucinations, near-perfect faithfulness, real citations) while presenting drug Y's clinical evidence as evidence about queried drug X. We term this deceptive grounding (DG): a failure invisible to faithfulness, hallucination, and citation checks because every claim is sourced from a real document, about the wrong entity. Using a controlled factorial benchmark across 13 models, we find DG rates spanning 8-87% at peak adversarial conditions. Medical and biomedical fine-tuned models reach up to 86.7%; domain specialization amplifies the failure rather than mitigating it. A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation. The two failure modes respond to the same trigger, taking different paths. Production measurement across 740 drug-disease pairs finds 7.8% overall DG in a deployed RAG system, rising to 13.6% for recently approved drugs. Entity-attribution verification (checking that cited evidence applies to the queried entity) detects DG at 97.0% precision and 98.7% DG recall (IPW-adjusted human gold standard); no existing framework implements it.

Comments24 pages, 7 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑