AI 中文总结
该研究针对阿尔茨海默病纵向预测中总体共形预测带对高风险亚组覆盖不足的问题,提出机制驱动的共形修正框架,在两个队列中实现了几乎所有高风险亚组的目标覆盖。
AI 中文摘要
阿尔茨海默病生物标志物的纵向预测越来越多地为临床决策提供依据,而预测结果只有在同时给出可信任程度时才有用。共形预测通过将任意预测器包裹在具有可交换性下有限样本覆盖保证的预测带中实现这一点。然而,标准的总体水平共形预测仅能保证边际覆盖,可能掩盖临床重要亚组内的严重覆盖不足。我们提出了一种通用的机制驱动框架,用于审计和修复此类亚组覆盖不足。在两个队列(ADNI、OASIS-3)、两个基础预测器以及涵盖遗传风险、人口统计学和临床严重程度的九个属性中,我们发现,尽管总体水平的预测带达到了名义边际覆盖,但在68个审计组合中有57个组合的高风险亚组覆盖不足。我们将这些失败归因于两种机制:(A)“稀有性”,即仅基于n名患者校准的组条件预测带的覆盖上限为k/(n+1);(B)“尾部厚重”,即总体水平的预测带对于尾部厚重的亚组过于狭窄,且额外数据无法缩小差距。覆盖不足不成比例地落在具有高遗传风险和疾病严重程度的患者身上(平均缺口6.1个百分点,95%置信区间[3.3, 8.9]),而人口统计学组平均保持在目标水平(0.0个百分点,置信区间[-1.9, 1.7])。我们为每种机制配备了相应的共形修正:针对稀有性的交叉共形池化、针对尾部厚重的亚组校准,以及当两种机制同时出现时的覆盖安全边际下限。这些修正共同恢复了两个队列和预测器中几乎每个高风险亚组的目标覆盖。
英文摘要
Longitudinal prediction of Alzheimer's disease biomarkers increasingly informs clinical decisions, and a forecast is only useful if it also reports how much to trust it. Conformal prediction supplies this by wrapping any forecaster in a prediction band with a finite-sample coverage guarantee under exchangeability. However, standard population-level conformal prediction guarantees only marginal coverage and may mask substantial under-coverage within clinically important subgroups. We introduce a general mechanism-driven framework for auditing and repairing such subgroup under-coverage. Across two cohorts (ADNI, OASIS-3), two base forecasters, and nine attributes spanning genetic risk, demographics, and clinical severity, we find that population-level bands under-cover high-risk subgroups in 57 of 68 audited combinations, despite achieving nominal marginal coverage. We trace these failures to two mechanisms: (A) \emph{rarity}, where a group-conditional band calibrated on only $n$ patients covers at most $k/(n+1)$; and (B) \emph{tail-heaviness}, where a population-wide band is too narrow for a heavy-tailed subgroup and additional data cannot close the gap. Under-coverage falls disproportionately on patients with high genetic risk and disease severity (6.1 pp mean deficit, 95\% CI [3.3, 8.9]), while demographic groups remain at the target level on average (0.0 pp, CI [$-1.9$, 1.7]). We pair each mechanism with a corresponding conformal correction: cross-conformal pooling for rarity, per-subgroup calibration for tail-heaviness, and a coverage-safe marginal floor when both arise. Together, these corrections restore target coverage for nearly every high-risk subgroup across both cohorts and forecasters.