arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考我们如何评估健康人工智能中的方法论进展

Rethinking How We Evaluate Methodological Progress in Health AI

Florent Pollet, Matthew McDermott

arXiv 2609.18134首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过统一框架重实现12种算法并在两个临床数据集上评估,发现算法比较结果跨任务族和数据集稳定转移,且新算法未必优于梯度提升树,表明方法论知识可能比假设的更需要少任务工程。

AI 中文摘要

电子健康记录(EHR)人工智能(AI)的方法论进展取决于我们判断哪些算法在何种条件下表现更优的能力。然而,此类进展被认为受到可重复性困难以及定义具有临床意义的评估任务难度的阻碍。我们通过在一个共享评估框架内重新实现12种历史及近期算法,并在两个临床数据集(MIMIC-IV和NWICU)上对其进行评估,来实证研究这些障碍。我们比较了两种互补的任务族:专家撰写的具有临床意义的任务,以及由随机采样的事件代码和预测时间范围定义生成的任务。我们探究相对算法比较是否能在任务族和数据集之间转移,残余的任务异质性是否包含有用的方法论结构,以及受控比较揭示了近十年来哪些进展。我们发现,总体成对比较在评估设置之间强烈转移,包括从随机生成的任务到具有临床意义的任务,以及跨数据集。同时,具有临床意义的任务表现出更大的任务-方法交互作用,这为任务属性有助于解释特定建模选择何时具有优势提供了初步证据。最后,较新的算法并未持续优于早期方法:梯度提升树在与现代、宽且稀疏的EHR表示结合时仍极具竞争力。总之,这些结果表明,有用的方法论知识可能比通常假设的需要更少的任务工程,同时强调了理解任务和方法之间剩余的结构化异质性的重要性。

英文摘要

Methodological progress in artificial intelligence (AI) for electronic health records (EHRs) depends on our ability to determine which algorithms work better, and under which conditions. However, such progress is thought to be hindered by difficulties in reproducibility and in defining clinically meaningful evaluation tasks. We empirically study these barriers by re-implementing 12 historical and recent algorithms within a shared evaluation framework and evaluating them on two clinical datasets, MIMIC-IV and NWICU. We compare two complementary task families: expert-authored clinically meaningful tasks and generated tasks defined from randomly sampled event codes and prediction horizons. We ask whether relative algorithms comparisons transfer across task families and datasets, whether residual task heterogeneity contains useful methodological structure, and what a controlled comparison reveals about progress over the last decade. We find that aggregate pairwise comparisons transfer strongly across evaluation settings, including from randomly generated tasks to clinically meaningful tasks and across datasets. At the same time, clinically meaningful tasks exhibit greater task-method interaction, providing preliminary evidence that task properties can help explain when particular modeling choices are advantageous. Finally, newer algorithms do not consistently outperform earlier approaches: gradient-boosted trees remain highly competitive when paired with a modern, wide and sparse representation of the EHR. Together, these results suggest that useful methodological knowledge may require less task engineering than commonly assumed, while highlighting the importance of understanding the structured heterogeneity that remains across tasks and methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑