arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.00635cs.LG

神经损失如何塑造VAE潜在变量

How Neural Losses Shape VAE Latents

  • Sapienza University of Rome(罗马大学萨皮恩扎分校)
  • Paradigma, Inc.(Paradigma公司)
  • Moises Systems, Inc.(Moises系统公司)
  • EPFL(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Giorgio Strano, Luca Cerovaz, Michele Mancusi, Tommaso Mencattini, Emanuele Rodolà

更新

AI总结:

本文研究感知损失和对抗损失等神经重建损失如何改变VAE的率失真问题,证明其减少潜在表示信息量并改变潜在空间几何结构,使表示更各向同性且不确定性分布更均匀。

AI中文摘要:

现代VAE很少使用标准$β$-VAE目标隐含的点态似然进行训练。在实践中,尽管缺乏对如何改变模型潜在动态的理解,点态重建常与感知损失和对抗损失结合。我们表明,重建损失的选择重塑了率失真问题本身,改变了潜在表示的信息内容和几何结构,这些变化可能仅从重建中无法察觉。首先,我们证明并实证验证,用神经项(如感知和对抗目标)增强点态重建会减少存储在潜在表示中的信息量。其次,我们展示神经重建损失系统地改变了潜在空间的几何结构:它们使表示更各向同性,并更均匀地将不确定性分布在潜在维度上,产生不同的后验方差分布。这些发现强调了率失真权衡并非理解VAE行为的全面视角,我们提出一种更机械的方法来研究失真度量的选择如何重塑优化问题。

英文摘要:

Modern VAEs are rarely trained with the pointwise likelihood implied by the standard $β$-VAE objective. In practice, pointwise reconstruction is often combined with perceptual and adversarial losses, despite a lack of understanding of how this changes the latent dynamics of the model. We show that the choice of reconstruction loss reshapes the rate-distortion problem itself, altering both the information content and the geometry of the learned latent space in ways that may be invisible from reconstructions alone. First, we prove and verify empirically that augmenting pointwise reconstruction with neural terms, such as perceptual and adversarial objectives, reduces the amount of information stored in the latent representations. Second, we show that neural reconstruction losses systematically change the geometry of the latent space: they make representations more isotropic and distribute uncertainty more evenly across latent dimensions, producing different posterior variance profiles. These findings highlight how the rate-distortion tradeoff is not a comprehensive lens to understand the behavior of VAEs, and we propose a more mechanistic approach to investigate how the choice of a distortion metric reshapes the optimization problem.

↑