arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17777stat.MEcs.AIq-bio.QM

信息集仿真:AI 派生 EHR 特征的因果证书

Information Set Emulation: Causal Certificates for AI Derived EHR Features

  • VRI
  • Department of Neurology, Juntendo University School of Medicine(顺天堂大学医学院神经科)

机构由 AI 辅助整理,请以论文原文为准。

Takes Fujita, Nobutaka Hattori

AI总结:

本文提出信息集仿真框架,为AI从电子健康记录提取的特征附加因果证书,通过类型化证据和纤维映射量化信息歧义,并给出识别与估计条件,明确点估计主张与兼容报告的适用场景。

AI中文摘要:

人工智能和大语言模型可以从电子健康记录(EHR)中恢复具有临床意义的特征,但预测有用性并不能确立其在因果推断中的可用性。我们引入了信息集仿真:在锁定的目标试验下,AI 类型化提升将来源证据、临床与记录时间、决策时可用性、表示版本、拟议的因果角色以及未解决的歧义附加到提取的特征上。因果证书记录了这些角色的可审计证据。具有未解决下游角色的特征被路由到兼容的报告或单独的分析中。类型化证据定义了与观测定律一致的因果世界的观测纤维。锁定的标量估计量将该纤维映射到一个兼容的图像,当图像非空且紧致时,其平方切比雪夫半径等于残差极小极大均方误差。这一经典恒等式提供了针对目标的信息歧义度量。其贡献在于与联合 EHR 观测图和可审计证书架构的集成。在明确的交换性、正性和干扰一致性条件下,我们给出了识别和交叉拟合增广逆概率加权估计,区分了经验目标和总体目标。EHR 压缩漂移恒等式分离了框架存在、治疗分配和结果观测的作用。人工模拟和一个普通法有限世界示例说明了估计失败和信息半径的缩减。合成第 0 阶段笔记展示了审计诊断;一个单独的角色特定分析展开说明了路由,并非精确的纤维半径。所有实验均为合成实验。该框架规定了重建信息何时能支持点估计主张,以及何时需要兼容的报告。

英文摘要:

AI and large language models can recover clinically meaningful features from electronic health records (EHRs), but predictive usefulness does not establish admissibility for causal inference. We introduce information set emulation: an AI typed lift attaches source evidence, clinical and recording times, decision-time availability, representation version, proposed causal roles, and unresolved ambiguity to extracted features under a locked target trial. Causal certificates record auditable evidence for those roles. Features with unresolved downstream roles are routed to compatible reporting or separate analyses. Typed evidence defines an observational fiber of causal worlds consistent with the observed law. The locked scalar estimand maps this fiber to a compatible image whose squared Chebyshev radius equals the residual minimax mean squared error when the image is nonempty and compact. This classical identity provides a target-specific measure of information ambiguity. The contribution is its integration with a joint EHR observation map and an auditable certificate architecture. Under explicit exchangeability, positivity, and nuisance-consistency conditions, we give identification and cross-fitted augmented inverse probability weighted estimation, distinguishing empirical and population targets. An EHR compression-drift identity separates the roles of frame presence, treatment assignment, and outcome observation. Artificial simulations and a common-law finite-world example illustrate estimation failures and information-radius reduction. Synthetic Phase 0 notes demonstrate audit diagnostics; a separate role-specific analysis spread illustrates routing and is not an exact fiber radius. All experiments are synthetic. The framework specifies when reconstructed information can support a point claim and when compatible reporting is required.

补充信息

↑