商业AI在肺栓塞及偶发性肺栓塞检测中的多站点真实世界性能
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
- Emory University School of Medicine(埃默里大学医学院)
- Yale University(耶鲁大学)
- Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究在17个设施的真实世界中评估了FDA批准的AI模型,发现其PE检测敏感性(86.8%)和iPE检测敏感性(73.5%)低于批准基准,但特异性更高,且对非急性及周围性栓塞性能下降,凸显了标准化上市后监测的必要性。
AI中文摘要:
肺栓塞(PE)是心血管死亡的主要原因,但FDA批准的AI检测模型在真实世界中的性能尚未得到充分表征。我们回顾性评估了来自单一商业平台(Aidoc Medical BriefCase)的两个FDA批准的AI算法,一个用于专用CT肺动脉造影(CTPA;n=30,678)上的PE分诊,另一个用于常规增强CT(n=37,191)上的偶发性PE(iPE)检测,覆盖了一个拥有17个设施的学术医疗系统。参考标准标签使用经过验证的LLM流程从放射学报告中提取(准确率97%,kappa=0.94)。PE模型实现了86.8%的敏感性和99.1%的特异性,敏感性从鞍状栓塞的99.3%下降到亚段PE的72.9%,从急性PE的89.7%下降到非急性PE的65.3%。iPE模型实现了73.5%的敏感性和99.8%的特异性。两个模型均表现出低于FDA批准基准的敏感性,而特异性超过批准基准,对周围性和非急性栓塞的性能下降反映了已知的人类阅片者局限性,并强调了需要对AI辅助医疗设备进行标准化的上市后监测。
英文摘要:
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.