AI 中文总结
本研究测量并解释胸部X光报告生成中的检索诱发幻觉,发现不匹配检索导致临床F1崩溃,并引入巧合重叠对照与病理一致性门控以缓解问题。
AI 中文摘要
检索增强生成是改进胸部X光报告的一种有吸引力的方式,因为来自相似既往研究的报告提供了通用视觉语言模型所缺乏的临床背景。我们证明同一机制是临床错误的可靠来源。在超过100项MIMIC-CXR研究中,使用冻结的LLaVA-1.5生成器和BioMedCLIP检索,CheXbert临床F1分数在相关检索下翻倍(从0.201升至0.402),而在临床不匹配检索下崩溃至0.043,仅为仅图像得分的五分之一;检索诱发幻觉率(仅计数可追溯至检索证据的不支持发现)从0.00升至0.76和0.98。为证明这不是证据集覆盖的伪影,我们引入巧合重叠对照,对仅图像生成进行评分,针对它们从未见过的证据,将偶然基线率定为0.18,比观察到的比率低四到五倍。生成器不仅获取发现,还转录文本:在相关检索下产生的95%的报告包含一个八词跨度,该跨度逐字出现在检索证据中但参考中缺失,而无检索时为0%。随后我们解释机制:正常胸部X光在28例中的24例中检索到至少一个异常先例,因为医学图像嵌入相似性主要由解剖和采集而非疾病存在主导。这有直接的设计后果。检索相似性不能预测危害(RIH率在相似性四分位数中为0.80/0.80/0.60/0.84),因此基于嵌入相似性的相关性门控无法工作;基于预测病理一致性的门控在无有用覆盖成本下移除61%的不支持证据。我们认为检索增强临床系统必须在检索失败而非仅检索成功下进行评估。
英文摘要
Retrieval-augmented generation is an attractive way to improve chest X-ray reporting, because reports from similar prior studies supply clinical context a general vision-language model lacks. We show the same mechanism is a reliable source of clinical error. Over 100 MIMIC-CXR studies with a frozen LLaVA-1.5 generator and BioMedCLIP retrieval, CheXbert clinical F1 doubles under relevant retrieval (0.201 to 0.402) and collapses to 0.043 under clinically mismatched retrieval, a fifth of the image-only score; the retrieval-induced hallucination rate, counting only unsupported findings traceable to retrieved evidence, rises from 0.00 to 0.76 and 0.98. To show this is not an artefact of evidence-set coverage, we introduce a coincidental-overlap control that scores image-only generations against evidence they never saw, placing the chance base rate at 0.18, four to five times below the observed rates. The generator does not merely acquire findings, it transcribes text: 95% of reports produced under relevant retrieval contain an eight-word span occurring verbatim in the retrieved evidence but absent from the reference, against 0% without retrieval. We then explain the mechanism: a normal chest X-ray retrieves at least one abnormal precedent in 24 of 28 cases, because medical image-embedding similarity is dominated by anatomy and acquisition rather than by the presence of disease. This has a direct design consequence. Retrieval similarity does not predict harm (RIH rates 0.80/0.80/0.60/0.84 across similarity quartiles), so relevance gates conditioned on embedding similarity cannot work; gating on predicted pathology agreement removes 61% of unsupported evidence at no cost to useful coverage. We argue that retrieval-augmented clinical systems must be evaluated under retrieval failure, not only under retrieval success.
Comments10 pages, 4 tables