发表机构
École Polytechnique Fédérale de Lausanne; Basque Center for Applied Mathematics (BCAM); University of the Basque Country (EHU)(洛桑联邦理工学院; 巴斯克应用数学中心; 巴斯克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过跨数据集评估发现,基于机器学习的咳嗽结核病分类器泛化能力差,采集设备等数据相关变异性是主因,强调其临床应用前需外部验证。
AI 中文摘要
咳嗽声学有望用于无创结核病(TB)筛查,但机器学习(ML)模型是捕捉了疾病相关声学特征还是数据采集的 artifacts(人工产物)仍未明确。我们在三个独立数据集上评估了经典机器学习和深度学习(DL)咳嗽型结核病分类器的跨数据集泛化能力。尽管数据集内表现中等(ROC-AUC最高达0.755±0.056),但两种 pipeline(流程)均无法泛化,外部表现常低于0.6,提示可能存在数据局限性。我们进一步观察到,音频表征按录音设备和数据集而非结核病状态组织,在CODA数据集中,预测的结核病概率与国家层面患病率相关,设备不匹配会降低迁移效果,而多样化设备训练可改善该问题。此外,临床变量基线泛化更一致(ROC-AUC为0.655-0.711),表明采集特定的变异性比群体偏移更易导致泛化性差。数据集内表现优异并不足够,咳嗽型结核病模型应用于临床前必须进行外部验证。
英文摘要
Cough acoustics are promising for non-invasive tuberculosis (TB) screening, yet whether machine learning (ML) models capture disease-related acoustics or artifacts of data collection remains unresolved. We evaluated the cross-dataset generalizability of classical ML and deep learning (DL) cough-based TB classifiers across three independent datasets. Despite moderate within-dataset performance (ROC-AUC up to $0.755 \pm 0.056$), both pipelines fail to generalize, with external performance frequently below 0.6, indicating a possible limitation of the data. We further observed audio representations are organized by recording device and dataset rather than TB status, predicted TB probability tracks country-level prevalence in CODA, and device mismatch degrades transfer while device-diverse training improves it. Additionally, a clinical-variable baseline generalizes more consistently (ROC-AUC $0.655 - 0.711$), indicating acquisition-specific variability is a stronger driver of poor generalizability than population shift. High within-dataset performance is not enough. External validation is essential before cough-based TB models are clinically ready.