arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

临床机器学习模型的变质测试:框架建议与初步研究

Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study

Jie JW Wu, Feiyu E, Bo Chen

arXiv 2607.22984首次发表:更新:

发表机构

Michigan Technological University(密歇根理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对临床机器学习模型,提出用变质测试评估其行为正确性,无需真实标签。设计12个候选变质关系目录及五层验证策略,在UCI心脏病数据集评估,发现临床模型有MT违反率,证明变质测试可补充传统指标评估临床预测模型。

AI 中文摘要

用于临床预测任务(如住院死亡率和败血症发作)的机器学习模型通常能获得较高的AUROC分数,但AUROC衡量的是排序性能而非临床敏感性。本文提出将变质测试(MT)应用于临床机器学习模型,以评估行为正确性,且无需单个预测的真实标签。利用MIMIC-III和MIMIC-IV数据集为三个ICU预测任务设计了12个候选变质关系目录,并提出五层验证策略。在UCI心脏病数据集上评估该方法,三个临床模型虽预测性能强,但在五个试点变质关系上的MT违反率为27%至87%。注入故障实验表明血压特征的符号否定错误AUROC未检测到,但MT违反率增加31至67个百分点。这表明变质测试是评估临床预测模型行为正确性的传统性能指标的有价值补充。

英文摘要

Machine learning models for clinical prediction tasks, such as in-hospital mortality and sepsis onset, routinely achieve high AUROC scores. However, AUROC measures ranking performance rather than clinical sensibility. A model may rank patients correctly overall while predicting a lower mortality risk when a patient's SOFA score worsens, contradicting established medical knowledge. This paper proposes applying metamorphic testing (MT) to clinical machine learning models to evaluate behavioral correctness without requiring ground-truth labels for individual predictions. We design a catalog of 12 candidate metamorphic relations (MRs) for three ICU prediction tasks using the MIMIC-III and MIMIC-IV datasets, with each MR grounded in an authoritative clinical guideline. We further propose a five-layer validation strategy to ensure that MRs are clinically sound before deployment. As a feasibility study, we evaluate the approach on the UCI Heart Disease dataset. Although the three clinical models achieve strong predictive performance (AUROC = 0.849-0.900), they exhibit MT violation rates ranging from 27% to 87% across five pilot MRs. An injected-fault experiment further shows that a sign-negation error in a blood pressure feature remains undetected by AUROC but increases the MT violation rate by 31-67 percentage points. These findings suggest that metamorphic testing provides a valuable complement to conventional performance metrics for assessing the behavioral correctness of clinical prediction models.

Comments4 pages, 1 figure, 4 tables. Accepted at the AIware 2026 arXiv Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑