arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同的预测,不同的伤害:患者世界模型的因果审计

Same Predictions, Different Harms: Causal Auditing of Patient World Models

Yicheng Qi, Xiyi Xiong

arXiv 2610.05198首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究审计患者世界模型在反事实伤害预测上的可靠性差距,提出两阶段共享响应SCM及敏感性边界方法,并在临床模拟器上验证,为安全信任提供协议。

AI 中文摘要

用于临床试验模拟的患者世界模型可以在转移核和分组风险上达成一致,但在切换治疗时受伤害的患者比例上可能存在分歧——这一反事实量对于干预感知推理至关重要。我们在一个两阶段共享响应结构方程模型(SCM)中审计了这一可靠性差距:一个分类的中间健康状态之后是常见的终端护理。在独立阶段下,尖锐的伤害区间在至多三个中间状态下具有闭式端点,并在四个状态处存在精确性边界。声明的依赖性和响应不匹配预算在放宽阶段独立性或完全中介时产生校准的外界;在一个对称的三状态模型中,整个敏感性前沿是尖锐的,$[0,\min\{1/2,1/3+(\rho+\delta)/2\}]$,并精确展示了预算如何消除超过仅端点界限的增益。两个八变量响应线性规划(LP)传播干预不确定性以进行有限样本审计。精确的见证者验证可达性。在公开的临床模拟器(EpiCare;脓毒症)上,原生配置显示出很少的阶段依赖性解决,并且没有额外的联合兼容性增益超过成对传输——这是对可靠性主张的诚实负面结果。所有实验均可本地复现;保证仍以所述因果模型为条件。结果提供了一个具体协议,用于决定何时患者世界模型可以安全地信任用于反事实伤害。

英文摘要

Patient world models used for clinical trial simulation can agree on transition kernels and arm-specific risks, yet disagree on the fraction of patients harmed by switching treatment---the counterfactual quantity that matters for intervention-aware reasoning. We audit this reliability gap in a two-stage shared-response SCM: a categorical intermediate health state is followed by common terminal care. Under independent stages, the sharp harm interval has closed-form endpoints for at most three intermediate states, with an exactness boundary at four states. Declared dependence and response-mismatch budgets yield calibrated outer bounds when stage independence or complete mediation is relaxed; in a symmetric three-state model the entire sensitivity frontier is sharp, $[0,\min\{1/2,1/3+(ρ+δ)/2\}]$, and shows exactly how budgets erase the gain over endpoint-only bounds. Two eight-variable response LPs propagate interventional uncertainty for finite-sample audits. Exact witnesses verify attainability. On public clinical simulators (EpiCare; sepsis), native configurations show little resolved stage dependence and no additional joint-compatibility gain over pairwise transport---honest negative results for reliability claims. All experiments are locally reproducible; guarantees remain conditional on the stated causal model. The results provide a concrete protocol for deciding when a patient world model is safe to trust for counterfactual harm.

Comments35 pages, 4 figures. Includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑