发表机构
Competence Center for Clinical Trials Bremen; University of Bremen; Leibniz Institute for Prevention Research and Epidemiology – BIPS(不来梅临床研究能力中心; 不来梅大学; 莱布尼茨预防研究与流行病学研究所——BIPS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LiD-GLM模型,利用可逆残差神经网络结合利普希茨约束,在平衡灵活性与可解释性的同时实现广义线性模型的非线性参数估计与分布假设修正,并通过事后正交化保证模型可识别性。
AI 中文摘要
将传统统计模型与神经网络(NN)组件结合为半结构化混合模型,是一种兼具传统可解释性与神经网络前所未有的灵活性的建模思路。为保留可解释性,通常需对神经网络组件加以限制以防止其主导模型,但现有对神经网络组件施加结构约束的方法会严重限制模型灵活性,而仅施加弱间接约束的方法则会失去有意义的可解释性。为此,本文提出的方法利用可逆残差神经网络(i-ResNets)为广义线性模型提供非线性参数估计与分布假设的灵活修正,同时保持建模分布对(原线性)预测因子的随机单调性。i-ResNets对应与恒等映射的可控偏差,通过约束其利普希茨常数,可严格限制并量化混合模型与传统模型的偏离程度,从而在不限制可学习非线性及交互效应结构的前提下,实现用户可指定的灵活性与可解释性权衡。此外,本文为模型开发了特定的固有解释技术,并通过适配的事后正交化确保模型可识别性。
英文摘要
The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit their models' flexibility; in contrast, methods that only enforce weak, indirect constraints lose meaningful interpretability. The method we propose therefore leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) predictor. The i-ResNets correspond to a controlled deviation from identity and by constraining their Lipschitz constant one can rigorously limit and quantify how far the hybrid model deviates from its traditional counterpart. This enables a user-specifiable compromise between flexibility and interpretability without limiting the structure of nonlinear and interaction effects that can be learned. Furthermore, we develop specific inherent interpretation techniques for our model and enforce model identifiability through an adapted post-hoc orthogonalization.
Comments25 pages, 15 figures