发表机构
Southern University of Science and Technology; University of California, Los Angeles; University of Hertfordshire(南方科技大学; 加州大学洛杉矶分校; 赫特福德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有限离散计数数据的过度离散性问题,本文提出Bernstein-二项模型,利用Bernstein分布构建灵活先验,并建立频率学派与贝叶斯推断框架,显著提升拟合精度与预测性能。
AI 中文摘要
有限离散计数数据在生物医学和社会科学等领域中普遍存在,且这些数据经常表现出过度离散性,使得二项分布不足以进行建模。尽管beta-二项分布通过为成功概率赋予主观的beta先验来缓解这一问题,但其对预先指定的参数形状的依赖可能导致先验信息的表示存在偏差,无法捕捉真实世界数据的复杂性和异质性特征。为解决这些局限性,本文提出了一种新型Bernstein-二项模型。通过利用Bernstein分布的均匀逼近性质,我们为二项分布中的成功概率构建了一个高度灵活且通用的先验框架,该框架能够动态适应复杂的数据结构,避免了传统单峰先验的主观性和限制。我们为频率学派和贝叶斯推断建立了全面的理论框架。此外,该模型被扩展到回归设置以考虑协变量效应。大量的模拟研究和真实世界数据应用证实,所提出的Bernstein-二项模型显著优于现有模型。在建模具有复杂潜在先验结构的过度离散计数时,它提供了更优的拟合精度、稳健性和样本外预测性能。
英文摘要
Finite discrete count data are ubiquitous in fields such as biomedicine and social sciences and these data frequently exhibit over-dispersion, rendering the binomial distribution inadequate for modeling. Although the beta-binomial distribution mitigates this by assigning a subjective beta prior to the success probability, its reliance on a pre-specified parametric shape may result in biased representations of prior information, failing to capture the complex and heterogeneous characteristics of real-world data. To address these limitations, this paper proposes a novel Bernstein-binomial model. By leveraging the uniform approximation properties of Bernstein distribution, we construct a highly flexible and general prior framework for the success probability in binomial distribition, which dynamically adapts to complex data structures, avoiding the subjectivity and restrictions associated with conventional unimodal priors. A comprehensive theoretical framework for both frequentist and Bayesian inferences is established. Furthermore, the model is extended to a regression setting to account for covariate effects. Extensive simulation studies and real-world data applications confirm that the proposed Bernstein-binomial model significantly outperforms existing models. It offers superior fitting accuracy, robustness, and out-of-sample predictive performance when modeling over-dispersed counts with complex latent prior structures.