arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06080cs.LGcs.AI

PhenoBench:深度表型人类队列能告诉我们什么

PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

  • Weizmann Institute of Science(魏茨曼科学研究所)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

Gal Sapir, Alon Diament, Adva Wolf, Doron Yaya-Stupp, Dikla Gelbard Solodkin, Dana Azouri, Anat Etzion-Fuchs, Guy Lutsker, Eran Segal, Hagai Rossman

AI总结:

PhenoBench是一个基于人类表型项目的可执行基准,定义了90项临床任务,通过标准化评估契约比较多种模型,揭示了预测价值的异质性及基础模型和语言模型的性能差距。

AI中文摘要:

深度表型队列结合了从秒到年时间尺度的临床、影像、分子和可穿戴设备观测数据。这种广度可以揭示哪些测量信息与哪些健康相关问题相关,但异质性分析之间无法直接比较。我们提出了PhenoBench,一个围绕人类表型项目构建的可执行基准,该项目已有超过13,000名参与者完成了初次访视。每个问题固定了目标、合格人群、时间点和允许的信息;其评估契约规定了数据划分、指标、基线和声明边界。该基准定义了涵盖15个领域和26种输入模态的90项临床任务。测量结果显示,预测价值因问题和表示而异,包括相对于匹配基线的保留性能的正向、接近零和负向变化。我们使用PhenoBench评估了新兴的表格基础模型,共进行了跨越52项任务的160项匹配回归比较。这些模型在总体上排名高于标准任务特定模型,但在每个单元格内对三个预训练模型取平均后,相对于岭回归的中位数改进仅为0.004 $R^2$(95%置信区间,0.002--0.006)。随后,我们使用相同的队列数据和评估契约评估了14个语言模型,共同覆盖了涵盖表型恢复、分类、随访预测和参与者排序的40项任务。在没有队列特定拟合的情况下,语言模型在某些任务上做出了有信息的预测,但显示出任务特定的能力差距、共同的规模失败,并且很少超过在同一字段上拟合的模型。PhenoBench将一个多模态纵向队列转变为一个版本化、可审计的评估系统,在该系统中,可以添加新问题、测量和模型,而无需重新定义现有比较。

英文摘要:

Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark that turns deep-phenotyping measurements into explicit questions and controlled comparisons of information sources and predictive models. It is built around the Human Phenotype Project, with more than 13,000 participants at the initial visit. Each question fixes the target, population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. PhenoBench defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Across 160 matched regression comparisons spanning 52 tasks, six pretrained tabular models ranked above the evaluated task-specific baselines, including XGBoost and CatBoost, under a fixed single-estimator protocol with bounded tuning. Giving each task equal weight, their mean advantage over ridge was 0.0103 $R^2$ (95% task-bootstrap interval, 0.0071-0.0136). We also evaluated 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks but showed task-specific capability gaps, shared failures of scale, and rarely surpassed task-specific ridge or logistic regression models fitted on the same input fields. PhenoBench provides a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.

补充信息

↑