发表机构
Erasmus University Medical Center; Erasmus University Rotterdam(伊拉斯姆斯大学医学中心; 伊拉斯姆斯大学鹿特丹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对ESM数据的缺失偏差与高维挑战,提出深度广义混合模型,结合变分自编码器与贝叶斯数据增广算法,在随机缺失数据下实现有效推断,经GrowIt!研究及模拟验证具应用潜力。
AI 中文摘要
体验抽样法(ESM)是一种纵向研究设计,参与者每日多次报告其想法、情绪状态与行为。本研究的动机源于GrowIt!应用程序收集的数据,该应用旨在调查新冠疫情期间青少年的日常情绪。当前分析ESM数据的方法面临诸多挑战:标准统计技术可能无法在高维场景中良好扩展,而机器学习方法则会因缺失数据引入的选择偏差产生有偏结果。在本研究的激励数据集里,青少年因此前强烈的负面情绪而退出,因此隐含的缺失数据属于标准机器学习方法无法处理的随机缺失(missing-at-random)类型。我们开发了一种新型神经网络架构,将混合效应模型推广至深度学习以克服这些挑战,该架构可通过固定效应与随机效应对数据的均值和相关结构进行半参数化且灵活的建模。在估计环节,我们采用了变分自编码器(variational auto-encoders)的适配版本与贝叶斯数据增广算法。通过该方法,模型可适配遵循通用分布的纵向结果、良好扩展至高维场景,并在数据为随机缺失时提供有效推断。我们将深度广义混合模型(Deep Generalised Mixed Model)应用于GrowIt!研究及多项模拟实验,结果显示该模型具有应用潜力,但因模型不稳定性存在次优性能。
英文摘要
The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the GrowIt! app, which was released to investigate daily emotions among adolescents during the COVID-19 pandemic. Current procedures to analyse ESM data face various challenges. While standard statistical techniques may not scale well to a high-dimensional setting, machine learning procedures can give biased results due to selection bias introduced by missingness. In our motivating dataset, adolescents dropped out due to previous strong feelings of negative emotions. Hence, the implied missing data are of the missing-at-random type that standard machine learning procedures cannot accommodate. We develop a novel neural network architecture that generalises mixed effects models to deep learning to overcome these challenges. It allows semi-parametric and flexible modelling of data's mean and correlation structure through fixed and random effects. For estimation, we use an adaptation of variational auto-encoders and a Bayesian data augmentation algorithm. Through this approach, the model can accommodate longitudinal outcomes following generic distributions, scale well to high-dimensional settings and provide valid inference when data are missing-at-random. We applied the Deep Generalised Mixed Model to the GrowIt! study and various simulations. The results show potential for the Deep Generalised Mixed Model, yet suboptimal performance due to model instability.