发表机构
Institute of Science Tokyo; Tsinghua University; Wuhan University; The University of Tokyo; Ant Group(东京科学研究所; 清华大学; 武汉大学; 东京大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Unite-Audio首次联合学习连续音频表示与潜在流匹配,通过生成目标塑造潜在空间并采用Flow-GRPO后训练,实现紧凑模型下的竞争性文本到音频生成性能。
AI 中文摘要
文本到音频(TTA)生成旨在合成真实反映自然语言描述的音频。大多数TTA系统采用两阶段潜在范式:先优化音频分词器以用于重建,然后将其冻结,再在所得潜在空间中训练生成模型。然而,面向重建的表示可能对生成而言并非最优,这促使人们考虑联合表示学习与生成学习。为此,我们引入Unite-Audio,据我们所知,这是首个联合学习连续音频表示和潜在流匹配以用于TTA的方法。通过将重建与自监督生成预测相结合,Unite-Audio使生成目标能够直接塑造潜在空间,而非将其视为固定的中间表示。我们进一步采用Flow-GRPO后训练来改善文本条件下的生成。实验表明,紧凑的潜在流模型取得了具有竞争力的TTA性能,而消融研究证实了联合学习音频表示和生成模型的益处。音频样本可在https://runwushi.github.io/Unite-Audio获取。
英文摘要
Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. However, reconstruction-oriented representations may be suboptimal for generation, motivating joint representation and generative learning. To this end, we introduce Unite-Audio, to our knowledge, is the first to jointly learn continuous audio representations and latent flow matching for TTA. By coupling reconstruction with self-supervised generative prediction, Unite-Audio allows the generative objective to directly shape the latent space rather than treating it as a fixed intermediate representation. We further employ Flow-GRPO post-training to improve text-conditioned generation. Experiments show competitive TTA performance with a compact latent flow model, while ablation studies confirm the benefit of jointly learning the audio representation and generative model. Audio samples are available at https://runwushi.github.io/Unite-Audio.