发表机构
University of Guelph(圭尔夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文强调临床机器学习中数据过滤决策的重要性,主张将其视为科学方法的一部分,并呼吁建立可解释、透明的预处理流程以理解其对数据分布和模型性能的影响。
AI 中文摘要
使用临床数据的机器学习(ML)研究通常在模型开发之前依赖预处理和过滤流程。这些流程中做出的过滤决策可以改变数据集的统计结构,并可能人为地降低或增加预测任务的复杂性。我们认为,过滤选择应被视为科学方法的一部分,而非常规的预处理步骤。我们进一步讨论了可解释且透明的预处理流程的必要性,这些流程使研究人员能够理解为何做出特定的过滤选择,以及这些选择如何影响最终的数据分布和模型性能。本工作的全部源代码可在GitHub上获取。
英文摘要
Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further discuss the need for explainable and transparent preprocessing pipelines that allow researchers to understand why specific filtering choices are made and how these choices affect the resulting data distribution and model performance. All of the source code for this work is available on GitHub.
CommentsPresented at AIMLSystems 2026 (Lecco, Italy)