发表机构
Elmore Family School of Electrical and Computer Engineering, Purdue University(普渡大学埃尔莫尔电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究变分自编码器潜在空间优化问题,提出软约束优化方法,包括基于熵的约束和权重过滤法,在dSprites和MNIST数据集上实验,提升了潜在变量激活分数、降低重建误差、减少潜在维度并加快收敛。
AI 中文摘要
变分自编码器(VAE)的有效性取决于其潜在空间的两个难以同时实现的属性:单个潜在变量的高编码能力以及这些变量的低维、解纠缠组织。削弱Kullback-Leibler正则化会提高能力但会降低解纠缠,而加强它会完全剔除潜在变量。我们将VAE训练公式化为一个软约束优化问题来解决这两个问题。首先,我们对单个潜在变量施加基于熵的约束(EC),表明潜在代码的熵上界了它携带的数据生成因子的互信息。其次,我们提出一种权重过滤方法,利用软约束的松弛在下游训练期间修剪低熵维度。在dSprites数据集上,EC使总体潜在变量激活分数比普通VAE提高43%-62%,在β-VAE变体中获得最高的FactorVAE分数(0.891对0.847),并将重建误差降低多达38%。在MNIST数据集上,权重过滤器将提供给下游分类器的潜在维度从十个减少到两个,同时保持准确率高于90%,比没有EC的相同过程少37%的轮次收敛。我们还发现低熵离散因子倾向于合并到单个潜在变量中,而高熵连续因子分布在几个潜在变量中。
英文摘要
The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity in the individual latent variables, and a low-dimensional, disentangled organization of those variables. Weakening the Kullback-Leibler regularization raises capacity but degrades disentanglement, while strengthening it prunes latent variables away entirely. We formulate VAE training as a soft-constrained optimization problem that addresses both. First, we impose an entropy-based constraint (EC) on individual latent variables, showing that the entropy of a latent code upper-bounds the mutual information it carries about the generative factors of the data. Second, we propose a weight-filter method that exploits the slack of the soft constraint to prune low-entropy dimensions during downstream training. On dSprites, the EC raises the aggregate latent-variable activation score by 43-62% over a vanilla VAE, attains the highest FactorVAE score among the \b{eta} \b{eta}-VAE variants (0.891 vs 0.847), and lowers reconstruction error by up to 38%. On MNIST, the weight filter reduces the latent dimensionality supplied to a downstream classifier from ten to two while holding accuracy above 90%, converging in 37% fewer epochs than the same procedure without the EC. We also find that low-entropy discrete factors tend to merge into a single latent variable, whereas high-entropy continuous factors are distributed across several.
Comments12 figures, 7 tables