arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于深度学习的脑部MRI重建安全性评估

Evaluating the Safety of Deep Learning-Based Brain MRI Reconstruction

Dat Tat Mai, Thai Viet Pham, Thu Nguyen Thi Dang, James Jin Kang

arXiv 2608.28714首次发表:更新:

发表机构

School of Science, Engineering & Technology, RMIT University Vietnam; School of Computing Technologies, RMIT University; School of Health and Biomedical Science, RMIT University(RMIT越南大学科学、工程与技术学院; RMIT大学计算技术学院; RMIT大学健康与生物医学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对深度学习脑部MRI重建的安全性问题,系统综述263项相关研究,发现现有评估无法捕捉病灶消除等故障,推导了安全评估需满足的5项要求。

AI 中文摘要

目的:深度学习可将脑部MRI扫描速度提升4至10倍,但模型可能会消除病灶或合成虚假组织——这类故障是PSNR、SSIM等像素平均指标无法捕捉的。本文综述当前评估实践是否能发现这一盲区。方法:遵循PRISMA 2020规范,无日期限制检索7个数据库,纳入1995至2026年间的263项研究,采用QUADAS-2及匹配工具评估,采用叙事综合法分析。分类源自标题、摘要及受控词汇;报告的患病率数值为下限。重复筛选一致性高(Fleiss kappa=0.877),评估一致性同样较高(0.788;可观测部分为0.390)。提取工作未经过审核。结果:263项研究中仅18项(6.8%)在相同数据上同时记录保真度指标与阅片者评估,核心替代指标未被测量。阅片者研究大多测量阅片者间一致性,结果较弱:fastMRI 2020的一致性达0.457和0.386(Kendall W),仅在SSIM出现差异时有所提升。在给定误差模型下,消除100 mm³的腔隙性梗死灶会使全局PSNR偏移0.03 dB。随研究总量增至5倍,阅片者评估比例从32%降至18%,后回升至21%。生成式模型最常与幻觉关联(39%),但接受阅片者评估的比例最低(11.3%);自监督模型达47%,且无任何阅片者评估。仅5%的研究发布代码并开展阅片者研究;无研究评估模型观察者;无命名数据集覆盖急性卒中或出血。结论:基于这些下限数据,当前评估实践无法验证诊断安全性。本文推导了面向安全的评估必须满足的5项要求。

英文摘要

Objective: Deep learning accelerates brain MRI four- to tenfold, but models can erase lesions or synthesize false tissue - failures pixel-averaged metrics like PSNR and SSIM miss. We review whether current evaluation practices detect this blind spot. Methods: Following PRISMA 2020, we searched seven databases without date limits, including 263 studies (1995-2026), appraised them using QUADAS-2 and matched instruments, and synthesized narratively. Categories were derived from titles, abstracts, and controlled vocabulary; reported prevalence figures represent floors. Duplicate screening achieved high agreement (Fleiss kappa = 0.877), as did appraisal (0.788; 0.390 where observable). Extraction is unaudited. Results: Only 18 of 263 studies (6.8%) recorded both a fidelity metric and reader assessment on identical data, leaving the central surrogate unmeasured. Reader studies mostly measured inter-reader agreement, which was weak: fastMRI 2020 concordance reached 0.457 and 0.386 (Kendall W), improving only where SSIM diverged. Erasing a 100 mm3 lacunar infarct shifts global PSNR by 0.03 dB under the stated error model. As the corpus grew fivefold, reader assessments dropped from 32% to 18%, recovering to 21%. Generative models - most associated with hallucination (39%) - were among the least reader-evaluated (11.3%), while self-supervised models reached 47% with zero reader evaluation. Only 5% released code and ran reader studies; none evaluated a model observer; no named dataset covered acute stroke or hemorrhage. Conclusions: On these floors, current evaluation practices cannot certify diagnostic safety. We derive five requirements safety-oriented evaluations must meet.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑