arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27224stat.APcs.LG

Psych-ECA:用于纵向精神病学中合成对照组的可复现半合成基准

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

Aakash Bhagat, Shashank Choudhary

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对精神病药物研发缺乏符合监管要求的合成对照组基准的问题,构建了Psych-ECA半合成基准,评估多种估计器性能,发现Scribe方法兼具准确性与校准性,逆强度校正可降低信息性采样偏差,成果可复现。

中文摘要 AI 辅助

外部对照组和合成对照组(ECAs)正逐步应用于精神病药物研发领域,但该领域缺乏能评估监管机构所关注特性的基准,这些特性不仅包括方法重建未处理轨迹的准确性,还涵盖其不确定性是否经过校准、是否对心理健康记录中常见的信息性观察时间(病情较重患者就诊更频繁)具有鲁棒性,以及会在“通过/不通过”试验决策中产生何种假阳性率。真实的精神病试验数据(如STAR-D及注册队列)需授权访问且缺乏真实反事实,因此,遵循因果推断领域已确立的半合成基准(IHDP、ACIC及PK-PD肿瘤生长模拟器),我们发布了Psych-ECA,这是一款可完全复现的纵向症状轨迹生成器,适用于抑郁症(PHQ-9量表)、焦虑症(HAM-A量表)和精神病(PANSS量表),具备已知的反事实对照组、信息性就诊安排及经验证的量表测量噪声。我们对八种估计器进行了基准测试,涵盖 carry-forward(向前结转)、汇总真实世界数据平均值、最近邻匹配、线性混合模型、梯度提升及Scribe轨迹桥接方法。研究得出三项发现:其一,轨迹方法与灵活机器学习方法实现了最佳反事实准确性(PHQ-9的RMSE约为2.3),优于横断面基线;其二,仅Scribe方法兼具准确性与校准性,其名义90%预测区间的经验覆盖率达93%-96%,而梯度提升为87%-88%,未校准的SDE模型则为62%-75%;其三,逆强度校正可降低信息性采样下的偏差,且Scribe的校准区间是唯一在信息性程度提升时仍能维持名义假阳性率的轨迹方法。我们发布了所有代码、数据生成脚本及随机种子,以支持完全可复现的评估工作。

英文摘要

External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions. Real psychiatric trial data (e.g. STAR-D and registry cohorts) require credentialed access and lack ground-truth counterfactuals, so, following established semi-synthetic benchmarks in causal inference (IHDP, ACIC, and the PK-PD tumor-growth simulator), we release Psych-ECA, a fully reproducible generator of longitudinal symptom trajectories for depression (PHQ-9), anxiety (HAM-A), and psychosis (PANSS) with known counterfactual control arms, informative visits, and validated-scale measurement noise. We benchmark eight estimators spanning carry-forward, pooled real-world-data averages, nearest-neighbour matching, linear mixed models, gradient boosting, and the Scribe trajectory-bridge method. Three findings emerge. First, trajectory and flexible machine learning methods achieve the best counterfactual accuracy (about 2.3 PHQ-9 RMSE), outperforming cross-sectional baselines. Second, only Scribe is both accurate and calibrated, achieving 93-96% empirical coverage of nominal 90% prediction intervals, compared with 87-88% for gradient boosting and 62-75% for uncalibrated SDE models. Third, inverse-intensity correction reduces bias under informative sampling, while Scribe's calibrated intervals are the only trajectory method that maintains nominal false-positive rates as informativeness increases. We release all code, data-generation scripts, and random seeds to enable fully reproducible evaluation.

发表机构

  • Sapien Labs(萨皮恩实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑