arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

患者世界模型中的重复测量泄漏、分布偏移与部分观测下的可靠性

Repeated-Measure Leakage, Distribution Shift, and Reliability under Partial Observation in Patient World Models

Arjun Subramanian

arXiv 2610.04778首次发表:更新:

发表机构

Massachusetts Institute of Technology; Computer Science and Artificial Intelligence Laboratory (CSAIL)(麻省理工学院; 计算机科学与人工智能实验室(CSAIL))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估患者世界模型在部分观测下的可靠性,发现重复测量泄漏和分布偏移显著影响预测性能,提出包含患者分离、重复测量控制、偏移和不确定性验证的评估栈。

AI 中文摘要

患者世界模型越来越多地被提出用于纵向预测、干预感知推理和临床试验模拟。因果或临床干预的有效性不同于预测泛化性和可靠性;在做出更强的主张之前,基础预测状态应能在患者间泛化,能应对现实中的偏移和缺失观测,并通过有意义的可靠性信号暴露失败。我们在一个刻意狭窄的设定中评估这些前提:基于PhysioNet GaitPDB的短时程数字生物标志预测,包含165名参与者、306条记录和51,129个上下文-未来对。使用持久性模型、岭回归、MLP、GRU、Transformer和一个紧凑的JEPA风格预测器,我们构建了一个评估阶梯,逐步移除原始时间重叠、同记录熟悉度和同患者熟悉度,然后测试未见患者的泛化性。对于GRU,NMSE从随机窗口划分下的0.1227上升到消除原始训练-测试重叠后的0.1393,并在患者留出法下达到0.1961。在54名具有重复记录的参与者中,暴露于同一患者的不同记录使GRU的NMSE从0.2177改善到0.1556,而记录排除的身份假设在参与者层面未得到支持。在参与者留出评估下,MLP和Transformer在统计上无显著差异。研究偏移、四倍长的预测间隔和部分观测进一步降低性能;在50%时间掩蔽下,Transformer的NMSE上升到0.611,而MC-dropout的预测方差下降。我们不声称构建了纵向或干预感知模拟器。相反,结果支持一个前提评估栈,包括患者分离、重复测量控制、偏移、缺失性和不确定性验证,然后才能信任更强的患者世界模型主张。

英文摘要

Patient world models are increasingly proposed for longitudinal prediction, intervention-aware reasoning, and clinical-trial simulation. Causal or clinical intervention validity is distinct from predictive generalization and reliability; before making stronger claims, the underlying predictive state should generalize across patients, survive realistic shifts and missing observations, and expose failure through meaningful reliability signals. We evaluate these prerequisites in a deliberately narrow setting: short-horizon digital-biomarker forecasting from PhysioNet GaitPDB, comprising 165 participants, 306 recordings, and 51,129 context-future pairs. Using persistence, ridge, MLP, GRU, Transformer, and a compact JEPA-style predictor, we build an evaluation ladder that progressively removes raw temporal overlap, same-recording familiarity, and same-patient familiarity before testing unseen-patient generalization. For GRU, NMSE rises from 0.1227 under random-window splitting to 0.1393 after eliminating raw train-test overlap and to 0.1961 under patient holdout. Among 54 participants with repeated recordings, exposure to a different recording from the same patient improves GRU NMSE from 0.2177 to 0.1556, while a recording-excluded identity hypothesis is not supported at the participant level. Under participant-held-out evaluation, MLP and Transformer are statistically indistinguishable. Study shift, a four-times-longer prediction gap, and partial observation further degrade performance; under 50% temporal masking, Transformer NMSE rises to 0.611 while MC-dropout predictive variance falls. We do not claim a longitudinal or intervention-aware simulator. Instead, the results support a prerequisite evaluation stack of patient separation, repeated-measure controls, shift, missingness, and uncertainty validation before stronger patient-world-model claims are trusted.

Comments13 pages, 7 figures. Selected for oral presentation at the NeurIPS 2026 Workshop on World Models for High-Stakes Health (WMHS) Replication Package: https://doi.org/10.17605/OSF.IO/AUZVR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑