发表机构
Gaston Berger University; LERSTAD(加斯顿贝格尔大学; LERSTAD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对离散双变量情形下的联合熵、条件熵和互信息,提出偏差校正估计量,推导其一阶偏差公式,并证明校正与未校正版本渐近等价,通过模拟和汽车保险应用验证了改进效果。
AI 中文摘要
本文研究了有限支撑离散随机变量的联合熵、条件熵和互信息的插件估计量的有限样本偏差。利用我们配套工作中发展的单变量约简框架,我们推导了显式的一阶偏差公式。对于支撑大小分别为r和s的一对(X,Y),偏差分别为-(rs-1)/(2n)、-r(s-1)/(2n)和(r-1)(s-1)/(2n)。然后,我们为这三个度量提出了偏差校正估计量,并证明了未校正和校正版本是渐近等价的,它们相差一个O(1/n)阶的确定性项;因此,它们具有相同的渐近分布。一项模拟研究验证了理论结果,包括专门的独立性研究以及与Wilks定理的关系。在汽车保险组合中的应用,其中X是驾驶员的年龄类别,Y是索赔严重程度类别,说明了偏差校正对风险分类带来的改进。
英文摘要
This paper studies the finite-sample bias of plug-in estimators for joint entropy, conditional entropy, and mutual information for finitely supported discrete random variables. Using the univariate reduction framework developed in our companion works, we derive explicit first-order bias formulas. For a pair (X,Y) with support sizes r and s, the biases are respectively -(rs-1)/(2n), -r(s-1)/(2n), and (r-1)(s-1)/(2n). We then propose bias-corrected estimators for the three measures and prove that the uncorrected and corrected versions are asymptotically equivalent, differing by a deterministic term of order O(1/n); consequently, they share the same asymptotic distribution. A simulation study validates the theoretical results, including a dedicated independence study and the relationship with Wilks' theorem. An application to a motor insurance portfolio, where X is the driver's age class and Y the claim severity class, illustrates the improvement brought by the bias correction for risk classification.
Comments20 pages, 6 figures, 3 tables. Keywords: Joint entropy; Conditional entropy; Mutual information; Bias correction; Univariate reduction; Asymptotic equivalence; Independence testing; Wilks' theorem; Motor insurance; Risk classification