一种基于阈值的可移植框架用于医学数据的可解释分类
Interpretable and Calibrated Classification of Clinical Data Using Supervised Feature Binarization
另 4 家 · 查看机构详情
- Worcester Polytechnic Institute(伍斯特理工学院)
- Centro de Vacunación e Investigación (CEVAXIN)(疫苗接种与研究中心(CEVAXIN))
- Universidad Tecnológica de Panamá(巴拿马技术大学)
- Massachusetts Institute of Technology(麻省理工学院)
- Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所)
- McGill University(麦吉尔大学)
- The Montreal Neurological Hospital-Institute(蒙特利尔神经学医院研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对医学数据分类中黑箱模型的问题,引入基于统计学的框架,用伯努利朴素贝叶斯模型并结合\(\chi^2\)引导的统计二值化方法,在多数据集上评估,性能与复杂模型相当,还提供可解释规则和校准风险估计。
中文摘要 AI 辅助
黑箱模型因缺乏可解释性和可重复性限制了人工智能在医学中的应用。我们引入了一个基于统计学的框架,使用伯努利朴素贝叶斯(BNB)模型提供完全可解释、基于规则的临床分类。该方法对连续变量应用监督式\(\chi^2\)引导的统计二值化,识别训练数据中与临床结果关联最大的阈值。在三个基准数据集上评估该方法,除了判别能力,还通过包括布里尔分数等方法评估概率可靠性。结果表明该可统计解释的框架能达到与更复杂模型相当的性能,还提供明确的临床决策规则和校准风险估计。
英文摘要
Black-box models limit the adoption of artificial intelligence in medicine because their predictions are difficult to interpret and reproduce. We present a statistically grounded framework for interpretable, rule-based clinical classification using the Bernoulli Naïve Bayes (BNB) model. Supervised chi-square-guided binarization converts continuous variables into binary indicators by selecting thresholds that maximize association with the clinical outcome within the training folds, which allows BNB to operate on continuous medical data without sacrificing transparency. On three benchmark datasets, Pima Indians Diabetes, Wisconsin Breast Cancer, and Heart Failure Prediction, the framework reached areas under the receiver operating characteristic curve of 0.800, 0.984, and 0.919, respectively. Probabilistic reliability was assessed with a leakage-safe cross-validated calibration analysis reporting Brier score and calibration intercept and slope, and post-hoc beta calibration improved probability calibration across datasets. These results indicate that an interpretable, statistically motivated framework can perform comparably to more complex models while providing explicit decision rules expressed in clinical units and calibrated risk estimates. A complete worked example further shows that model inference can be reproduced from a printed reference table using only basic arithmetic, without software or proprietary tools, supporting trustworthy and auditable use of artificial intelligence in clinical settings.