在预后回归模型合成场景中区分病例混合与背景异质性
Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings
AI总结:
该研究提出方法区分预后回归模型合成中的病例混合与背景异质性,经COPD试验验证,可指导选择联合或站点特定回归模型。
AI中文摘要:
预后回归模型常需合成来自多个站点的数据,这些站点既可以是多中心研究内部的、联邦学习场景下的,也可以是基于个体参与者数据的元分析中的站点;此处的站点指任意数据源,如医院、登记处、试验或研究,无需为物理中心。分析人员需决定是采用单一回归模型代表所有站点,还是需要站点特定的模型。现有系数水平tau²等指标可量化异质性,但无法区分其来源。我们聚焦于诊断系数异质性反映的是病例混合还是站点特定背景效应:当线性回归项近似不同协变量分布人群中的多变量非线性关系时,会产生病例混合异质性;当可比患者在不同站点需要不同的回归关系时,则会产生背景异质性。我们通过在降维空间中拟合站点特定的局部回归,并将平滑后的系数表面划分为跨站点参考和站点特定偏差来实现这一目标;我们采用autoencoder(自动编码器)和自定义损失在局部预后关系周围构建潜在空间,随后将该划分投影到结局尺度以推导观测级和站点级汇总指标。我们在包含两个站点的COPD试验上验证了该方法:在三个主要潜在斜率坐标中,系数表面变化主要为背景性的;推导得到的观测级结局尺度方差划分以病例混合为主,而其站点间聚合则集中于背景差异而非病例混合偏移。我们还采用置换站点阴性对照评估当站点标签无信号时,背景汇总指标是否会出现,该诊断区分可用于指导应评估联合回归模型还是站点特定回归模型。
英文摘要:
Prognostic regression models often synthesize data from multiple sites, whether within a multi-site study, across federated settings, or in individual participant data meta-analysis. Here, a site is any data source, such as a hospital, registry, trial, or study, and need not be a physical center. Analysts must then decide whether one regression model represents all sites or whether site-specific models are needed. Established measures such as coefficient-level tau^2 quantify heterogeneity but do not distinguish its source. We focus on diagnosing whether coefficient heterogeneity reflects case-mix or site-specific context effects. Case-mix heterogeneity can arise when linear regression terms approximate multivariable non-linear relationships in populations with different covariate distributions. Contextual heterogeneity arises when comparable patients require different regression relationships across sites. We do this by fitting site-specific local regressions in a dimension-reduced space and partitioning the smoothed coefficient surfaces into a cross-site reference and site-specific deviations. An autoencoder and custom loss structure the latent space around local prognostic relationships. We then project this partition onto the outcome scale to derive observation- and site-level summaries. We demonstrate the approach on a COPD trial with two sites. In the three leading latent slope coordinates, coefficient-surface variation was predominantly contextual. The derived observation-level outcome-scale variance partition was case-mix-leading, whereas its between-site aggregation was concentrated in contextual differences rather than case-mix shifts. A permuted-site negative control assesses whether the contextual summary can arise when site labels carry no signal. This diagnostic distinction can inform whether joint or site-specific regression models should be evaluated.