arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越模型:数据过滤在临床机器学习中的关键作用

Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning

Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin, Andrew Hamilton-Wright

arXiv 2610.06640首次发表:更新:

发表机构

University of Guelph(圭尔夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文强调临床机器学习中数据过滤决策的重要性,主张将其视为科学方法的一部分,并呼吁建立可解释、透明的预处理流程以理解其对数据分布和模型性能的影响。

AI 中文摘要

使用临床数据的机器学习(ML)研究通常在模型开发之前依赖预处理和过滤流程。这些流程中做出的过滤决策可以改变数据集的统计结构,并可能人为地降低或增加预测任务的复杂性。我们认为,过滤选择应被视为科学方法的一部分,而非常规的预处理步骤。我们进一步讨论了可解释且透明的预处理流程的必要性,这些流程使研究人员能够理解为何做出特定的过滤选择,以及这些选择如何影响最终的数据分布和模型性能。本工作的全部源代码可在GitHub上获取。

英文摘要

Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further discuss the need for explainable and transparent preprocessing pipelines that allow researchers to understand why specific filtering choices are made and how these choices affect the resulting data distribution and model performance. All of the source code for this work is available on GitHub.

CommentsPresented at AIMLSystems 2026 (Lecco, Italy)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑