arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KMGen:一种基于技能的合成个体患者数据生成方法

KMGen: A Skill-based Approach for Synthetic Individual Patient Data Generation

Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun

arXiv 2608.22618首次发表:更新:

发表机构

Mayo Clinic Alix School of Medicine(梅奥诊所阿利克斯医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KMGen是首个端到端框架,可全自动提取KM曲线并生成合成患者不良事件轨迹,在多肿瘤试验中精度达标,相关流水线已开源。

AI 中文摘要

来自临床试验的个体患者数据(IPD)是生存建模、荟萃分析和安全性研究的基础,但IPD很少被公开。现有工作仅解决了这一缺口的一半:从已发表的图表中重建Kaplan-Meier(KM)曲线——通常需要手动数字化或人在回路的修正——却未提供构成患者记录另一半的不良事件(AE)流的生成机制。我们推出KMGen,这是首个端到端框架,它(i)全自动提取KM曲线,精度可与人工引导工具媲美,(ii)从公开试验注册记录中生成合成的每位患者AE轨迹。提取阶段是一个完全自动化的智能体流水线——智能体生成代码以提取KM曲线的每一步——在包含干净、边缘情况和对抗条件的32个图表基准上,实现了0.0151的平均积分绝对误差(IAE)。IPD生成阶段将患者原型提取与统计采样解耦:大型语言模型(LLM)将试验记录提炼为特定治疗组的统计数据、不良事件、患者人口统计学特征和风险乘数。机械采样器通过临床原型、与经验KM曲线的自举秩相关耦合(精确保留边际生存分布),以及带诱导/维持拆分的基于周期的AE调度来生成患者事件。在三个保留的肿瘤学试验中,队列规模相差一个数量级,且每个试验进行30次独立再生,KMGen的平均积分KM绝对差Δ_KM≤0.051,6个人口统计学条目中有5个性别/ECOG的JSD≤0.013,且在单一固定参数集下,通过精确MedDRA术语恢复了≥71%的前15种AE。该流水线作为开源代码发布在this https URL。

英文摘要

Individual patient data (IPD) from clinical trials is the substrate for survival modeling, meta-analysis, and safety research, yet IPD is rarely released. Prior work has addressed only half of this gap: reconstructing Kaplan-Meier (KM) curves from published plots -- typically requiring manual digitization or human-in-the-loop correction -- while offering no mechanism for generating the adverse-event (AE) streams that constitute the other half of a patient record. We introduce KMGen, the first end-to-end framework that (i) fully automates KM curve extraction at accuracy competitive with human-guided tools, and (ii) generates synthetic per-patient AE trajectories from public trial registry records. The extraction stage is a fully automated agentic pipeline -- an agent generates code to extract each step in the KM curve -- achieving a mean Integrated Absolute Error (IAE) of 0.0151 on a 32-plot benchmark spanning clean, edge-case, and adversarial conditions. The IPD generation stage decouples patient archetype extraction from statistical sampling: an LLM distills the trial record into arm-specific statistics, adverse events, patient demographics, and risk multipliers. A mechanistic sampler generates patient events via clinical archetypes, bootstrap rank-correlation coupling to the empirical KM curve (preserving the marginal survival distribution exactly), and cycle-based AE scheduling with an induction/maintenance split. Across three held-out oncology trials spanning an order of magnitude in cohort size and 30 independent regenerations per trial, KMGen achieves mean integrated KM absolute difference $Δ_{\text{KM}}\,{\leq}\,0.051$, sex/ECOG JSD ${\leq}\,0.013$ on 5 of 6 demographic slots, and recovers ${\geq}\,71\%$ of the top-15 AEs by exact MedDRA term under a single fixed parameter set. The pipeline is released as open source at https://github.com/chufangao/kmgen.

Journal refProceedings of Machine Learning Research 340 (2026) 1-60

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑