Copula 活动子空间 I:用于降阶非高斯密度估计的得分-协方差方法
Copula Active Subspaces I: A Score-Covariance Method for Reduced-Order Non-Gaussian Density Estimation
浏览论文内容
中文总结 AI 辅助
提出 Copula 活动子空间方法,通过 copula 得分协方差矩阵的特征分解实现非高斯噪声密度的降阶表示,在贝叶斯推断中显著降低 KL 散度。
中文摘要 AI 辅助
在具有非高斯观测噪声的贝叶斯推断问题中,后验分布的准确性取决于噪声密度的准确性,而基于梯度的采样器需要该密度及其梯度能够在逐点上求值,无论来自显式表达式还是代码,且无需内部求解。我们提出 Copula 活动子空间(CAS)来表示这种噪声密度。分量秩变换将噪声法则的相依性隔离在其 copula 中,秩-r 约简仅保留相依性变化的方向。这些方向是 copula 得分协方差矩阵 $\boldsymbol{C}:= \mathrm{Cov}_{\pi_{\boldsymbol{Z}}}(\nabla\log c^{Z})$ 的主特征向量,这使该约简成为 copula 活动子空间。由于当坐标独立时 $\boldsymbol{C}$ 为零,这些是相依性方向,数据协方差未必能识别这些方向。由此构造得出高斯参考 KL 散度界,其显式常数为 $\tfrac{1}{2}$,在所有秩-r 约简中由该特征空间精确最小化;一个仅从样本即可计算的约简遗留误差诊断;以及通过 Hermite 得分匹配,得到闭式约简对数密度和梯度,截断阶数和阶段 2 正则化常数在验证样本上选择。该约简将 d 维密度估计问题替换为 r 维问题。在 d=20 的噪声法则和具有该噪声的贝叶斯推断问题中,与高斯 copula、边际乘积和 PCA 子空间基线相比,CAS 将噪声 KL 散度降低五倍以上,后验 KL 散度降低七倍以上,并在另外两个 d=20 示例中将噪声 KL 散度降低约 3.5 倍和 2.7 倍。
英文摘要
In Bayesian inference problems with non-Gaussian observation noise, the posterior is only as accurate as the noise density, and gradient-based samplers need that density and its gradient evaluable pointwise, whether from an explicit expression or from code, and without an inner solve. We propose Copula Active Subspaces (CAS) to represent this noise density. A componentwise rank transform isolates the noise law's dependence in its copula, and a rank-$r$ reduction keeps only the directions along which that dependence varies. These directions are the leading eigenvectors of the copula score covariance $\boldsymbol{C} := \mathrm{Cov}_{π_{\boldsymbol{Z}}}(\nabla\log c^{Z})$, which is what makes the reduction a copula active subspace. Because $\boldsymbol{C}$ vanishes when the coordinates are independent, these are directions of dependence, which the covariance of the data need not identify. From this construction follow a Gaussian-reference KL divergence bound with the explicit constant $\tfrac{1}{2}$, minimized over all rank-$r$ reductions by exactly this eigenspace; a diagnostic for the error the reduction leaves behind, computable from the samples alone; and, from Hermite score matching, a reduced log-density and gradient in closed form, with the truncation orders and the Stage-2 regularization constants chosen on validation samples. The reduction replaces a $d$-dimensional density estimation problem by an $r$-dimensional one. On a $d=20$ noise law and a Bayesian inference problem with that noise, CAS lowers noise KL divergence more than fivefold and posterior KL divergence more than sevenfold against Gaussian-copula, product-of-marginals, and PCA-subspace baselines, and lowers noise KL divergence by factors of about $3.5$ and $2.7$ on two further $d=20$ examples.
发表机构
- Colorado State University(科罗拉多州立大学)
机构由 AI 辅助整理,请以论文原文为准。