arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12360cs.CYcs.LG

监管批准并不足够:FDA clearance医疗设备中可信赖AI报告的缺口

Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

Ahmed M Salih, Oliver Díaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir

首次发表
浏览论文内容

中文总结 AI 辅助

该研究分析2021-2025年FDA获批的519份AI/ML医疗设备报告,发现可信赖AI报告存在显著缺口,获批年份或临床领域不影响报告透明度,仅监管批准不足以证明AI可信赖性。

中文摘要 AI 辅助

背景:支持AI/ML的医疗设备在不断发展的监管框架下越来越多地部署在医疗保健领域。随着这些系统越来越融入临床决策,人们越来越期望它们展现可信赖AI的关键维度,以支持临床医生、患者和公众的信任。公开可用的监管文件是否提供了足够的证据来独立评估已获批AI系统的可信赖性,仍不清楚。方法:我们分析了2021年至2025年间发布的FDA支持AI/ML的医疗设备总结报告。这些报告经过自动关键词筛选,随后进行多阶段人工共识审查,以确定六个FUTURE-AI原则(公平性、通用性、可追溯性、可用性、鲁棒性和可解释性)的记录证据。进行了描述性、时间性和临床领域分析。多变量逻辑回归评估了获批年份或临床领域是否预测更高的报告透明度,透明度定义为报告了三个或更多原则的证据。结果:在筛选的1105份FDA总结报告中,519份被纳入。可信赖AI报告有限且不均衡。近四分之一(24.7%)的报告未提供任何原则的证据,没有一份记录了所有六个原则的证据。鲁棒性是报告最频繁的(57.6%),而可追溯性(8.3%)和可解释性(3.5%)是最明显的缺口。获批年份(OR 1.02,95% CI 0.88-1.19)或临床领域(OR 0.73,95% CI 0.46-1.15)均未预测更高的报告透明度。解释:FDA文件中存在大量且持续的可信赖AI报告缺口。仅监管批准不应被视为可信赖性的替代指标。需要在AI生命周期内进行标准化、可审计的报告,以支持医疗AI的独立评估和负责任采用。

英文摘要

Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether year of clearance or clinical domain predicted higher reporting transparency, defined as evidence reported for three or more principles. Results: Of 1,105 FDA summary reports screened, 519 were included. Trustworthy AI reporting was limited and uneven. Nearly one quarter (24.7%) provided no evidence for any principle, and none documented evidence across all six. Robustness was most frequently reported (57.6%), while Traceability (8.3%) and Explainability (3.5%) were the most pronounced gaps. Neither year of clearance (OR 1.02, 95% CI 0.88-1.19) nor clinical domain (OR 0.73, 95% CI 0.46-1.15) predicted higher reporting transparency. Interpretation: Substantial, persistent trustworthy AI reporting gaps exist in FDA documentation. Regulatory approval alone should not be considered a proxy for trustworthiness. Standardised, audit-ready reporting across the AI lifecycle is needed to support independent assessment and responsible adoption of healthcare AI.

发表机构

  • University of Leicester(莱斯特大学)
  • British Heart Foundation (BHF) Leicester Centre of Research Excellence(英国心脏基金会莱斯特卓越研究中心)
  • Universitat de Barcelona(巴塞罗那大学)
  • Chalmers University of Technology(查尔姆斯理工大学)
  • Universidad de Sevilla(塞维利亚大学)
  • Queen Mary University of London(伦敦玛丽女王大学)
  • Institució Catalana de Recerca i Estudis Avançats (ICREA)(加泰罗尼亚高级研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑