arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05691cs.SD

AudioGAR:弥合潜在音频生成模型中的重建与生成差距

AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models

Xianghong Fang, Geeyang Tay, Wentao Ma, Tim G. J. Rudner, Dehan Kong

首次发表
浏览论文内容

中文总结 AI 辅助

AudioGAR通过扰动编码器潜在表示并利用冻结的潜在扩散模型构建中间潜在表示,仅微调解码器,以1.5%的训练数据和0.26%的成本显著提升潜在音频生成模型的端到端生成性能。

中文摘要 AI 辅助

潜在音频生成模型通常分两个阶段训练:首先学习音频编解码器,然后训练潜在生成模型。这种分解导致解码器训练与生成之间的不匹配:编解码器的解码器在编码器产生的潜在表示上训练,但在推理时却部署在生成器产生的潜在表示上。在多种数据集和潜在生成模型上,我们在FD和FAD指标下都观察到明显的重建-生成差距,这表明强大的重建质量并不一定能转化为强大的端到端生成质量。一个自然的补救措施是在生成产生的潜在表示上调整解码器,但生成的潜在表示与源音频缺乏对应关系,因此无法直接提供用于解码器微调的配对监督。我们提出AudioGAR,通过扰动编码器潜在表示并通过冻结的潜在扩散模型去噪来构建中间潜在表示。这些潜在表示形成从重建到生成的轨迹,其中低噪声潜在表示保留源对应关系并支持配对解码器微调。我们仅在这些潜在表示上微调解码器,同时保持编码器和潜在生成模型冻结。当应用于AudioX时,AudioGAR显著提升了生成性能。它仅需要原始训练音频时长的1.5%和原始训练成本的0.26%。

英文摘要

Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5\% of the original training audio hours and 0.26\% of the original training cost.

发表机构

  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑