arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格数据基础模型评估的比较框架:医疗保健案例研究

A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare

Majid Lotfian Delouee, Sjors G. J. G. In 't Veld, Martijn C. Schut

arXiv 2609.22154首次发表:更新:

AI 中文总结

本文提出一个比较评估框架,从六个临床维度对表格基础模型进行评分排名,并应用于医疗用例,提供45个模型分类,以指导模型选择。

AI 中文摘要

表格数据是临床实践中最常见的格式,涵盖实验室结果、用药记录、诊断代码和患者人口统计学信息。随着用于表格数据的基础模型在数量和种类上的增长,一个实际问题变得更加难以回答:对于给定任务,临床医生或数据科学家究竟应该选择哪个模型,以及为什么?现有综述对这些模型的功能进行了编目,但未能提供一种结构化的方式,使其能够针对真实应用的具体需求进行比较。我们引入了\system{},一个比较评估框架,该框架在六个临床相关的维度上对表格基础模型(TFMs)进行评分和排名:模型对新数据集的泛化能力、保护患者隐私的有效性、达到良好性能所需的数据量、随数据集和特征空间增长时的扩展性、预测对临床医生的可解释性,以及在不同患者亚组中的公平性表现。每个维度被分解为可测量的子组件,并且子组件组可选地组合成补充的复合分数,称为超级指标,提供模型在多个维度上同时表现的诊断视图。为了展示该框架的实际运作方式,我们将其应用于两个医疗保健用例:铁缺乏症筛查和心力衰竭预测,展示了相同的一组指标如何根据每个临床背景中最关键的因素导致不同的模型排名。我们还提供了按底层架构组织的45个TFM的分类,作为研究人员和从业者导航这一快速扩展领域的参考。

英文摘要

Tabular data is the most common format in clinical practice, encompassing laboratory results, medication records, diagnostic codes, and patient demographics. As foundation models for tabular data have grown in number and variety, a practical question has become harder to answer: which model should a clinician or data scientist actually choose for a given task, and why? Existing surveys catalogue what these models can do, but they stop short of providing a structured way to compare them against the specific demands of a real application. We introduce \system{}, a comparative evaluation framework that scores and ranks tabular foundation models (TFMs) across six clinically meaningful dimensions: how well a model generalizes to new datasets, how effectively it protects patient privacy, how much data it needs to perform well, how it scales with growing datasets and feature spaces, how interpretable its predictions are to clinicians, and how fairly it performs across patient subgroups. Each dimension is broken down into measurable sub-components, and groups of sub-components can optionally be combined into supplementary compound scores, called super-metrics, that provide a diagnostic view of how a model performs across several dimensions simultaneously. To show how the framework works in practice, we apply it to two healthcare use cases, screening for iron deficiency and predicting heart failure, demonstrating how the same set of metrics leads to different model rankings depending on what matters most in each clinical context. We also provide a taxonomy of 45 TFMs organized by their underlying architecture, which serves as a reference for researchers and practitioners looking to navigate this rapidly expanding field.

Comments32 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑