发表机构
Harvard Medical School; Clalit Research Institute; Clalit Health Services(哈佛医学院; 克拉利特研究所; 克拉利特健康服务)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究利用蒸馏自540万患者私人健康记录的数据,训练图注意力网络评分工具,使智能体在不暴露数据的情况下发现血液生物标志物,外部验证中显著提升AUC,支持隐私保护假设生成。
AI 中文摘要
常规全血细胞计数(CBC)可能产生新的生物标志物,但评估候选标志物所需的私人记录无法与擅长发现的前沿语言模型智能体共享。我们将Clalit Health Services中超过540万名患者的数据面板中的证据蒸馏为一个已发布的评分工具:针对13种免疫介导疾病中的每一种,在数据边界内训练了一个图注意力网络,用于预测候选CBC表达式的病例对照AUC,且仅发布训练后的权重。该工具在真实世界数据基础上支撑智能体的提出-评分-优化循环,而不暴露任何患者数据。在外部验证中,智能体发现的表达式相比其文献种子起点,中位数提高了4.18个AUC百分点;在三个独立队列中,对三个前沿研究工具的候选进行重新排序,在大多数比较中优于其首选,收益因队列而异。已发布的评分器支持隐私保护的生物标志物假设生成。
英文摘要
Routine complete blood counts (CBCs) could yield new biomarkers, but the private records needed to evaluate candidates cannot be shared with frontier language model agents that excel at discovery. We distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool: for each of 13 immune-mediated diseases, a graph attention network was trained inside the data boundary to predict the case-control AUC of candidate CBC expressions, and only the trained weights were released. The tool grounds an agent's propose-score-refine loop in real-world data without exposing any patient data. In external validation, agent-discovered expressions improved on their literature-seeded starting points by a median of 4.18 AUC percentage points, and across three independent cohorts, reranking the candidates of three frontier research tools improved on their first choices in most comparisons, with gains that varied by cohort. The released scorer supports privacy-preserving biomarker hypothesis generation.