发表机构
School of Computer Science and Electronic Engineering, University of Surrey; Meta Superintelligence Labs; Department of Informatics, King’s College London(萨里大学计算机科学与电子工程学院; 元宇宙超级智能实验室; 伦敦国王学院信息学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出FlowSep2,一种结合Self-Flow与Diffusion Transformer的文本条件流匹配生成模型,用于语言查询音频源分离,在多个基准上达到SOTA性能,可有效分离重叠声源。
AI 中文摘要
语言查询音频源分离(LASS)旨在根据自然语言描述从混合音频中提取目标声源,为音频源分离提供了灵活且可扩展的接口。然而,现有的大多数LASS方法依赖于判别式的基于掩码的模型,这类模型从输入混合音频中估计掩码,往往会过度抑制目标声音或无法完全分离它们,尤其是在复杂声学场景中多个声音事件强烈重叠时。本研究提出FlowSep2,一种用于LASS的文本条件流匹配生成模型。FlowSep2不直接预测分离掩码,而是在潜在空间中根据混合音频表示和文本查询的条件,从高斯噪声中学习生成目标声源表示。具体而言,我们采用带有Diffusion Transformer骨干网络的整流流匹配,进一步将自监督流匹配范式Self-Flow融入LASS框架。通过在生成目标下鼓励语义结构化的潜在表示,Self-Flow提升了模型根据文本查询分离目标声源的能力。在多个LASS基准上的实验表明,FlowSep2实现了最先进的性能,并在声音事件重叠的具有挑战性场景中展现出增强的声音分离效果。
英文摘要
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-based models, which estimate masks from the input mixture. These methods often over-suppress target sounds or fail to fully separate them, especially when multiple sound events strongly overlap in complex acoustic scenes. In this work, we propose FlowSep2, a text-conditioned flow-matching generative model for LASS. Instead of directly predicting a separation mask, FlowSep2 learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query. Specifically, we employ rectified flow matching with a Diffusion Transformer backbone. We further incorporate Self-Flow, a self-supervised flow-matching paradigm, into our LASS framework. By encouraging semantically structured latent representations under the generative objective, Self-Flow improves the model's ability to separate target sources according to text queries. Experiments on multiple LASS benchmarks show that FlowSep2 achieves state-of-the-art performance and demonstrates enhanced sound separation results in challenging scenarios with overlapping sound events.
CommentsSubmission to IEEE/ACM Transactions on Audio, Speech, and Language Processing