arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经验变分自编码器

Empirical Variational Autoencoder

Kaede Shiohara

arXiv 2610.06545首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出经验变分自编码器(EVA),通过经验学习自回归潜在先验替代标准高斯约束,缓解先验-后验分布差距,实现高保真序列生成,在图像和声音合成上达到与扩散模型相当的质量且推理更快。

AI 中文摘要

我们提出了经验变分自编码器(Empirical Variational Autoencoder, EVA),这是一个用于连续值(即非向量量化)序列的通用生成框架。EVA基于变分自编码器(VAE)的证据下界,但通过训练数据经验性地学习自回归潜在先验,这仅需在VAE之上增加一个额外的线性层即可实现。通过用自预测先验替代传统的标准高斯约束,EVA显著缓解了传统VAE中通常观察到的先验与后验之间的潜在分布差距,并为序列数据生成带来了高保真的祖先采样。在图像和声音合成上的大量实验表明,尽管EVA的推理时间快得多,但其生成质量可与自回归扩散基线相媲美。

英文摘要

We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the self-predicted priors, EVA significantly alleviates the latent distribution gap between prior and posterior which is typically observed in conventional VAEs, and leads to high-fidelity ancestral sampling for sequential data generation. Extensive experiments on image and sound synthesis demonstrate that EVA achieves competitive generation quality with autoregressive diffusion baselines despite its much faster inference time.

CommentsProject page: https://mapooon.github.io/EVAPage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑