AI 中文总结
该研究针对概率密度函数两阶段估计的偏差问题,提出了基于贝叶斯希尔伯特空间、中心对数比变换和ZB样条的惩罚最大似然框架,经模拟和地球化学数据验证了其有效性。
AI 中文摘要
概率密度函数通常通过初步平滑或聚合程序进行估计,例如直方图或核密度估计,之后再进行后续的函数表示和函数数据分析。这种两阶段方法可能会导致额外的近似偏差,并削弱观测数据与潜在分布结构之间的直接联系。本文提出了一种惩罚最大似然框架,用于在贝叶斯希尔伯特空间框架内从原始观测值直接估计概率密度函数,同时保留密度的成分几何结构。该方法基于中心对数比(clr)变换,这是贝叶斯希尔伯特空间与积分为零的标准平方可积勒贝格空间之间的等距同构,可实现高效的样条表示。经clr变换的密度使用ZB样条基函数表示,其平滑性通过对样条系数施加二次惩罚来控制。所提出的框架针对单变量和双变量密度进行了开发;对于双变量密度,它自然地将正交分解纳入独立部分和交互部分,以及相应的几何边际。其性能在涉及多个复杂场景的模拟研究中进行了评估,并与核平滑进行了比较。最后,使用经验地球化学数据说明了该框架的适用性。
英文摘要
Probability density functions are commonly estimated through preliminary smoothing or aggregation procedures, e.g., histograms or kernel density estimation, before subsequent functional representation and functional data analyses. Such a two-stage approach can lead to additional approximation bias and weaken the direct connection between the observed data and the underlying distributional structure. In this paper, we propose a penalized maximum likelihood framework for direct estimation of probability density functions from raw observations within the framework of Bayes Hilbert spaces while preserving the compositional geometry of densities. The methodology is based on the centered log-ratio (clr) transformation, an isometric isomorphism between Bayes Hilbert spaces and the standard Lebesgue space of square integrable functions with zero integral, enabling efficient spline representations. The clr transformed densities are represented using ZB-spline basis functions, while their smoothness is controlled through quadratic penalties imposed on the spline coefficients. The proposed framework is developed for univariate and bivariate densities. In the latter case, it naturally incorporates the orthogonal decomposition into independent and interactive parts together with the corresponding geometric marginals. The performance is evaluated in a simulation study involving multiple complex scenarios and compared with kernel smoothing. Finally, the applicability of the framework is illustrated using empirical geochemical data.