发表机构
Rutgers University(罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对潜变量生成模型中先验与聚合后验不匹配的问题,提出聚合后验预测检验(APPC)方法,理论上证明其渐近校准性,实验表明聚合后验采样能改善生成质量。
AI 中文摘要
潜变量生成模型通常使用简单的潜变量先验进行拟合,但从这些先验中抽取的样本往往无法产生逼真的数据。这种失败源于先验与聚合后验(即由拟合模型和数据诱导出的潜变量分布)之间的不匹配。这种不匹配通常被视为先验设定错误并应被替换的证据。或者,在现代生成模型中,越来越多地采用一种两阶段策略:首先拟合模型,其次估计聚合后验(van den Oord 等人,2017;Rombach 等人,2022)。然后通过从该聚合后验而非先验中采样来获得合成数据。为了检验此类过程,我们引入了聚合后验预测检验(APPC)。理论上,我们建立了APPC渐近校准的充分条件。对于概率主成分分析,我们表明,当普遍因素允许恢复信号空间时,APPC在潜先验设定错误的情况下仍能保持校准。使用变分自编码器的实验表明,与高斯先验采样相比,聚合后验采样改善了重尾和聚类数据的生成,同时其性能与使用更灵活潜先验的模型相当。
英文摘要
Latent variable generative models are commonly fit using simple priors over latent variables, but draws from these priors often fail to produce realistic data. This failure is due to a mismatch between the prior and the aggregated posterior, the distribution of latent variables induced by the fitted model and the data. This mismatch is often viewed as evidence that the prior is misspecified and should be replaced. Alternatively, in modern generative models, a two-stage strategy is increasingly used where first, the model is fit, and second, the aggregated posterior is estimated (van den Oord et al.,2017; Rombach et al., 2022.). Synthetic data are then obtained by sampling from this aggregated posterior instead of the prior. To check such procedures, we introduce the aggregated posterior predictive check (APPC). Theoretically, we establish sufficient conditions under which the APPC is asymptotically calibrated. For probabilistic principal component analysis, we show that the APPC can remain calibrated under a misspecified latent prior when pervasive factors permit recovery of the signal space. Experiments with variational autoencoders show that aggregated posterior sampling improves generation for heavy-tailed and clustered data relative to Gaussian prior sampling while performing comparably to models with more flexible latent priors.