arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有非参数边缘的贝叶斯异质copula混合模型:身体体能数据中的一致性、可识别性与尾部不对称性

Bayesian heterogeneous copula mixtures with nonparametric margins: consistency, identifiability and tail asymmetry in physical fitness data

Yujian Liu

arXiv 2610.08317首次发表:更新:

发表机构

Shanghai University of Sport(上海体育大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出两阶段贝叶斯伪后验方法估计copula混合模型,证明其一致性与可识别性,并在体能数据中发现尾部不对称模式,优于高斯copula。

AI 中文摘要

Clayton、Gumbel、Frank和高斯copula的有限混合可以描述上尾与下尾之间不同的相依性。我们研究了一种两阶段贝叶斯分析,其中用秩或核估计替代边缘分布,并在由此产生的伪观测值上评估copula似然。只要对数密度具有对数边界包络,该伪后验对copula密度就具有强一致性。两个阶段可以使用相同的数据,边缘可以在观测到的层内进行标准化,并且既不需要光滑性也不需要可识别性。该包络适用于任意固定维数下高斯、Student、Clayton、Gumbel和Frank copula的有限混合;因此,尾部相依系数和条件尾部概率可以被一致地估计。我们进一步证明了Clayton、Gumbel、Frank和高斯copula是联合有限线性独立的,这使得混合权重和分量可识别且可一致估计。在模拟中,伪后验在大样本下与多重启动的最大伪似然相匹配。在小样本和接近独立(每个分量都接近独立copula)的情况下,它更加稳定,此时尾部系数在权重之前就被学习到,并且边缘Metropolis采样器的效率最高可达数据增强的两倍。两个身体体能数据集显示出镜像不对称性。在8772名大学生中,短跑和跳跃表现主要在高端耦合。在国民健康调查的5336名成年人中,低握力和低日常活动聚集在一起,并随年龄增长而加剧,而高值则不然。高斯copula在留出数据中遗漏了这两种模式。

英文摘要

Finite mixtures of Clayton, Gumbel, Frank and Gaussian copulas can describe dependence that differs between the upper and lower tails. We study a two-stage Bayesian analysis in which ranks or kernel estimates replace the margins and the copula likelihood is evaluated at the resulting pseudo-observations. This pseudo-posterior is strongly consistent for the copula density whenever the log density admits a logarithmic boundary envelope. Both stages may use the same data, margins may be standardized within observed strata, and neither smoothness nor identifiability is required. The envelope holds for finite mixtures of Gaussian, Student, Clayton, Gumbel and Frank copulas in any fixed dimension; tail-dependence coefficients and conditional tail probabilities are therefore consistently estimated. We further prove that Clayton, Gumbel, Frank and Gaussian copulas are jointly finitely linearly independent, which makes mixture weights and components identifiable and consistently estimated. In simulations the pseudo-posterior matches multi-start maximum pseudo-likelihood in large samples. It is more stable in small samples and near independence (every component close to the independence copula), where tail coefficients are learned long before the weights and a marginal Metropolis sampler is up to twice as efficient as data augmentation. Two physical fitness datasets show mirror-image asymmetries. Among 8772 university students, sprint and jump performance are coupled mainly at the top. Among 5336 adults in a national health survey, low grip strength and low daily activity cluster together, increasingly with age, whereas high values do not. A Gaussian copula misses both patterns in held-out data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑