发表机构
Humboldt-Universität zu Berlin; Hertie Institute for AI in Brain Health; Universität Tübingen(柏林洪堡大学; 赫蒂脑健康人工智能研究所; 蒂宾根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对深度学习中遗漏变量偏差问题,提出基于广义加性模型的控制变量方法,通过交叉拟合岭惩罚重拟合预训练网络最后一层,在模拟与真实神经影像数据上验证了方法的有效性。
AI 中文摘要
控制变量在统计建模中被广泛用于解释已知混淆变量的遗漏变量偏差,但它们在深度学习中却鲜有研究,这一情况令人惊讶,因为当人口统计变量等可从图像中推断出的协变量与结果相关时,深度学习模型会将这些协变量编码到预测中,这种遗漏变量偏差被称为“捷径学习”。虽然现有的许多混淆控制或公平性方法试图限制此类协变量与模型预测的相关性,但我们表明这无法纠正遗漏变量偏差。因此,我们提出一种基于广义加性模型的深度学习模型控制变量方法,用于建模模型输入和协变量的效应。由于灵活的加性模型可能存在共曲线性,我们引入一种估计过程,通过使用带岭惩罚的交叉拟合,重新拟合预训练网络的最后一层以纳入协变量效应。我们展示了如何使这些效应相对于协变量正交化以排除其介导效应,以及如何在协变量分布上边缘化模型预测以控制其效应。这产生了无偏、可解释的预测,并可根据科学或公平性目标灵活建模所需效应。我们使用模拟图像验证了我们的方法,证明其能一致地估计真实效应,而现有方法要么需要更多数据,要么无法恢复真实效应。我们将该方法应用于存在实验诱导混淆的真实神经影像数据,其预测性能恢复至接近在无混淆数据上训练的模型水平,代码可在指定URL获取。
英文摘要
Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, and demonstrate consistent estimation of true effects. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data. Code is available at https://github.com/mpff/cocodeel.
Comments28 pages, 14 figures. Code at https://github.com/mpff/cocodeel