发表机构
University of Louisville; Case Western Reserve University; Mass General Brigham, Harvard Medical School; University of Missouri(路易斯维尔大学; 凯斯西储大学; 麻省总医院布里格姆医院,哈佛医学院; 密苏里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出贝叶斯混合模型,结合协变量的逻辑回归识别纵向数据中三类异质性来源,经模拟和SWAN研究的DHEAS数据验证,可准确识别异质性并有效估计固定效应参数。
AI 中文摘要
纵向数据中常存在异质性,部分观测组的均值和方差与其余组存在显著差异,部分观测组在少量测量点还会出现异常值。采用假设同质性的标准混合效应模型,会导致残差方差被高估,估计效率低下。本研究识别并处理纵向数据中三类异质性来源:不兼容的均值轨迹、增大的残差方差,以及单个测量点的异常值。所提出的贝叶斯混合模型纳入了针对这些特征的异质性二元指示器,通过使用协变量的逻辑回归进行建模。我们采用马尔可夫链蒙特卡罗方法进行统计推断,并实施模型选择以评估各类异质成分的纳入情况。模拟结果表明,该模型可准确识别异质性并生成固定效应参数的有效估计。我们还使用来自SWAN研究的DHEAS激素数据对所提方法进行了验证。
英文摘要
We often observe heterogeneity in longitudinal data, where the mean and variance for certain profiles meaningfully differ from the rest. Some profiles may also exhibit outliers at a limited number of measurements. Using a standard mixed effects model, which assumes homogeneity, can lead to overestimating the residual variance and inefficient estimation. In this work, we identify and account for three sources of heterogeneity in longitudinal data: incompatible mean trajectories, increased residual variance, and outliers at individual measurements. Our Bayesian mixture model incorporates binary indicators of heterogeneity for each of these features, modeled through logistic regression using covariates. We perform statistical inference using Markov chain Monte Carlo and implement model selection to evaluate the inclusion of various heterogeneous components. Simulations demonstrate that our model can accurately identify heterogeneity and produce efficient estimates of the fixed effects parameters. We further validate our approach using the DHEAS hormone data from the SWAN study.