发表机构
VRI; Department of Neurology, Juntendo University School of Medicine(VRI; 顺天堂大学医学部神经科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成协变量在序贯实验中角色不明导致因果问题改变的问题,提出因果类型规范与估计目标锁,通过标准化分解和聚类级正交估计量,确保在生成表示下估计目标保持有效。
AI 中文摘要
由笔记、对话、图像和可穿戴设备数据流生成的AI协变量,当其角色未被明确指定时,可能改变因果问题。一个生成的特征可能代表处理版本、行动前状态、历史、设计变量、中介变量、结果代理、观测过程或并发事件;这些角色不可互换。我们为序贯实验制定了一套因果类型规范:一个版本化表示映射、一个因果角色分类器、一个声明状态过滤器,以及一个估计目标锁。该锁在生成协变量进入分析之前固定一个标准化的近端效应。在审计正确性和标准识别假设下,可允许的角色分配保持该估计目标。我们将已建立的压缩偏差的条件协方差刻画应用于用生成表示替换设计相关状态。一个标准化的分解将压缩漂移、条件法则漂移和标准化漂移分开。进一步的结果涵盖中介调整、行动后泄漏、标记-干预混淆、结果引导发现以及状态测量误差。在重复会话和缺失结果下,聚类级正交估计量区分经验目标与超总体目标。模拟表明,当细化保留设计相关信息时有助于改进,而设计擦除、泄漏和同数据标记选择可能导致偏差或覆盖不足。该框架在生成表示进行确证性推断之前,将因果语义和声明状态置于首位。
英文摘要
AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.
Comments29 pages. Ancillary files include simulation code, seeds, and replicate-level results