arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

极值广义线性模型:高维情形下的估计与推断

Generalized Linear Models for Extremes: Estimation and Inference in High Dimensions

Liujun Chen, Chen Zhou

arXiv 2608.16137首次发表:更新:

AI 中文总结

本研究提出一种极端尾部回归模型,通过协变量缩放尾部尺度且不改变形状,采用Bregman散度匹配与ℓ₁惩罚实现高维估计,推导去偏估计量的渐近性质并应用于保险理赔数据。

AI 中文摘要

我们提出一种针对响应变量极端尾部的回归模型,其中协变量对尾部进行尺度缩放但不改变其形状。单个依赖协变量的函数即可刻画整个条件尾部,这与针对预先指定水平分位数的极端分位数回归形成对比。尾部形状本身不受限制:重尾、轻尾和短尾响应均被同一框架覆盖。我们通过连接函数和协变量的线性组合来设定该函数,这符合广义线性模型的思路。在估计阶段,我们在最大观测值构成的局部区域内,基于Bregman散度将参数化设定与潜在尾部函数进行匹配。由此得到的损失函数是凸的,且通过ℓ₁惩罚项可使协变量数量超过有效样本量。尾部局部化使得渐近理论有别于经典惩罚广义线性模型的渐近理论,统计分析仅使用由随机阈值选出的尾部观测值,因此这些观测值具有依赖性。我们推导了惩罚估计量的收敛速度,并提出一种具有渐近正态性的去偏估计量,可得到单个系数的置信区间。其渐近方差由得分的协方差决定,在尾部局部化下该协方差与海森矩阵不同,必须单独估计。我们将该方法应用于汽车保险理赔数据。

英文摘要

We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level. The tail shape itself is left unrestricted: heavy-, light- and short-tailed responses are covered by the same framework. We specify the function through a link function and a linear combination of the covariates, which is in the spirit of a generalized linear model. In estimation, we match the parametric specification to the underlying tail function under a Bregman divergence, over a region localized at the largest observations. The resulting loss is convex, and an $\ell_1$-penalty allows the number of covariates to exceed the effective sample size. The tail localization makes the asymptotic theory deviate from that for classical penalized generalized linear models. Only the tail observations selected by a random threshold are used in the statistical analysis, making them dependent. We derive the convergence rate of the penalized estimator and propose a debiased estimator that is asymptotically normal, yielding confidence intervals for individual coefficients. Its asymptotic variance is determined by the covariance of the score, which under tail localization differs from the Hessian and must be estimated separately. We apply the method to automobile insurance claims data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑