arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25234stat.AP

用于稳健高效拟合与验证对数位置尺度损失模型的分位数与对数分位数最小二乘

Quantile and Log-Quantile Least Squares for Robust-Efficient Fitting and Validation of Log-Location-Scale Loss Models

Mohammed Adjieteh, Vytaras Brazauskas

AI总结:

本文提出四种分位数最小二乘估计量,构建对数位置尺度损失模型的估计与验证框架,通过模拟与真实数据验证其性能,可适配不同规模样本。

AI中文摘要:

保险及其他类型损失的多种模型均属于对数位置尺度族的特例,其中对数正态分布与帕累托I型(Pareto-I)分布是最典型的例子,后者也是风险管理中常带来挑战的无限均值模型的主要实例。本文利用两个渐近定理——独立同分布随机变量样本分位数的联合正态性与德尔塔方法(delta method),构建用于对数位置尺度分布估计与验证的非线性及线性回归框架。在这些回归框架内,提出四种同等稳健的估计量:普通分位数最小二乘(oQLS)、广义分位数最小二乘(gQLS)、对数普通分位数最小二乘(log-oQLS)与对数广义分位数最小二乘(log-gQLS)。对于对数位置尺度损失模型,分位数的对数变换可精确近似非线性最小二乘解,且能得到更准确的估计量。此外,对数线性回归框架为研究估计量性质及设计基于残差的拟合优度检验提供了便捷方式。而且,log-oQLS与log-gQLS具有显式公式,可轻松用于中等规模(样本量n=10³)、大规模(n=10⁴)及超大规模(n>10⁶)样本。本文通过模拟数据集与真实数据集,展示了这些估计量、异常值标记规则及拟合优度检验的计算性能与统计性能。

英文摘要:

\begin{quote} {\bf\em Abstract\/}. ~A variety of models for insurance and other types of losses are special cases of the {\em log-location-scale\/} family, with the lognormal and Pareto-$I$ distributions being the most prominent examples. The latter also serves as a primary example of infinite-mean models that often present challenges in risk management. In this paper, we utilize two {\em asymptotic\/} theorems -- the joint normality of sample quantiles (of {\em i.i.d.\/} random variables) and the delta method -- to construct nonlinear and linear regression frameworks for estimation and validation of log-location-scale distributions. Within these regression frameworks, four equally robust estimators -- ordinary and generalized quantile (oQLS and gQLS) and log-quantile (log-oQLS and log-gQLS) least squares -- are proposed. For log-location-scale loss models, the logarithmic transformation of quantiles approximates the nonlinear least squares solution {\em exactly\/} and yields more accurate estimators. Also, the log-linear regression framework facilitates a convenient way to study the estimators' properties and to design a residuals-based goodness-of-fit test. Moreover, log-oQLS and log-gQLS have explicit formulas and can be easily computed for medium- ($n=10^3$), large- ($n=10^4$), and very large-size ($n > 10^6$) samples. Computational and statistical performances of the estimators, outlier-labeling rules, and the goodness-of-fit test are illustrated using simulated and real datasets. \vspace{2mm} {\bf\em Keywords\/}. ~Goodness-of-Fit; Outliers; Quantiles; Relative Efficiency; Robustness. \end{quote}

↑