发表机构
University of Illinois Urbana-Champaign; Mayo Clinic(伊利诺伊大学厄巴纳-香槟分校; 梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究临床数据插补问题,提出SafeImpute框架,通过构建事件图、利用双关系GNN和自适应融合学习插补,并结合共形选择控制不可接受误差,实验证明该方法在插补准确性和可靠误差控制上优于基线。
AI 中文摘要
临床护理通常依赖关键实验室指标,但实际患者就诊稀疏且检查安排不规则,导致大量数据缺失。许多插补方法虽提高平均准确率,但对高风险下游使用中哪些插补值足够可靠指导有限。本文研究可靠临床插补,旨在准确插补并选择性发布可靠结果,统计控制不可接受误差。提出SafeImpute框架,构建事件图捕捉患者内时间轨迹和患者间临床相似性,用双关系GNN和自适应融合学习插补,由辅助掩码重建目标正则化。通过转换代理风险分数为共形p值并应用Benjamini-Hochberg程序控制错误发现率。实验表明SafeImpute在插补准确性和可靠误差控制方面表现出色,优于多种基线方法。
英文摘要
Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness. While many imputation methods improve average accuracy, they provide limited guidance on which imputed values are reliable enough for high-stakes downstream use. In this work, we study reliable clinical imputation, aiming to produce accurate imputations while selectively releasing the reliable results, with statistical control over clinically unacceptable errors. To achieve this goal, we propose SafeImpute, a reliable imputation framework for irregular and sparse clinical longitudinal records. SafeImpute constructs an event graph that captures both intra-patient temporal trajectories and inter-patient clinical similarity, and learns imputations with a two-relation GNN and adaptive fusion, regularized by an auxiliary masked reconstruction objective. For reliability guarantees, SafeImpute converts a proxy risk score into conformal p-values and applies the Benjamini--Hochberg procedure to control the false discovery rate (FDR) of unacceptable errors among released imputations at a user-specified tolerance. Experiments on our Mayo Clinic data, the public MIMIC-III and MIMIC-IV datasets show that SafeImpute achieves strong imputation accuracy while providing reliable error control, outperforming diverse baselines in both standard imputation evaluation and FDR-controlled selective-release evaluation.
CommentsAccepted at KDD 2026. Author accepted manuscript