发表机构
Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出解耦潜在流匹配框架,通过VAE与Flow2GAN相关技术实现少步联合人声-伴奏分离,在减少采样预算时可提升分离性能与感知质量。
AI 中文摘要
生成建模为建模混合条件下的源分布提供了灵活方式,但迭代扩散和流匹配模型对于长音乐信号计算成本较高。本文通过潜在流匹配研究联合人声-伴奏分离:预训练变分自编码器(VAE)将混合信号与源信号映射到紧凑潜在空间,流匹配模型联合生成人声和伴奏的潜在表示。该框架通过分离编码器和速度解码器将语义分离与声学速度预测解耦,为降低采样成本,还借鉴Flow2GAN应用潜在对抗后训练以实现少步生成。实验表明,在减少采样预算下,潜在对抗优化可提升感知质量和分离指标。
英文摘要
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.