arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12477cs.LG

基于反事实结果专家标注的治疗诱导标签不确定性下的学习:以神经预后为例

Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication

Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对治疗决策导致部分患者结局不确定的临床预后问题,提出结合确定与不确定病例标签的预测模型及分类型评估框架,揭示两类病例评估的权衡关系。

中文摘要 AI 辅助

临床预测模型的开发通常假设每位患者的目标结局都能被清晰观测。但当治疗决策导致临床相关结局永久不可观测时,该假设不再成立。作为该问题的案例研究,我们使用包含2497名患者的队列开展心脏骤停后神经预后研究,其中1429名患者的结局因治疗决策变得不确定。这些结局不确定的患者由独立临床专家进行评估,专家提供了他们对患者本会出现的反事实结局的推测,我们将这类患者称为不确定病例;其余患者的临床相关结局可被观测,我们将其称为确定病例。我们提出了一种预测模型评估框架,该框架明确将评估分为确定病例和不确定病例两类,由于两类病例的可用目标标签存在差异,无法轻易以统一方式评估这两类病例。随后,我们提出了一种简单的预测模型,该模型同时使用确定病例和不确定病例的目标标签,允许我们在两类病例间进行权衡。在我们提出的神经模型与一系列表格基线模型中,具有相似确定病例AUROC的模型,在确定病例Brier评分和对不确定病例的概率估计方面却存在显著差异。我们提出的模型若要更好地匹配不确定病例的目标标签,通常需要以牺牲确定病例的准确性为代价,这凸显了标准评估所掩盖的明确权衡关系。这些结果表明,当治疗决策决定临床有意义的结局是否可观测时,传统评估指标会忽略对预后支持最为重要的患者群体中存在的重要失效模式。

英文摘要

Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. As a case study of this problem, we consider post-cardiac-arrest neurological prognostication using a cohort of 2,497 patients, including 1,429 patients whose outcomes were rendered indeterminate by treatment decisions. These patients with indeterminate outcomes were reviewed by independent clinical experts, who provided their guesses of counterfactual outcomes about what would have happened to the patients. We refer to these patients as uncertain cases. We also have patients for whom we observe their clinically relevant outcomes; we refer to these patients as certain cases. We propose a framework for evaluating prediction models that explicitly splits the evaluation between certain and uncertain cases. Here, we cannot easily evaluate both types of cases in a uniform manner as the available target labels differ. We then propose a simple prediction model that uses target labels from both certain and uncertain cases in a manner that allows us to trade off between them. Across the proposed neural model and a collection of tabular baselines, models with similar certain-case AUROC can nevertheless differ substantially in both certain-case Brier score and their probability estimates for uncertain cases. Improving alignment with target labels of uncertain cases for our proposed model generally comes at the cost of worse accuracy on certain cases, highlighting an explicit tradeoff that standard evaluation conceals. These results show that when treatment decisions determine whether clinically meaningful outcomes remain observable, conventional evaluation metrics can miss important failure modes in the very patients for whom prognostic support matters most.

补充信息

↑