通过MIMO模型扩展实现具有混合一致性的迭代音频分离
Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
浏览论文内容
中文总结 AI 辅助
本文提出一种通用框架,通过将源分离模型扩展为多输入多输出(MIMO)配置,实现稳定有效的迭代音频分离,同时保持混合一致性,并在最先进模型上取得显著性能提升。
中文摘要 AI 辅助
本文提出一个通用框架,通过将源分离模型扩展为多输入多输出(MIMO)配置,实现稳定且有效的具有混合一致性的迭代音频分离。在音频分离领域,混合一致性是许多需要目标源精确相位和音色信息的应用所必需的重要性质。虽然诸如扩散模型之类的迭代方法在语音增强或用户引导的目标源分离任务中取得了感知上优越的结果,但大多数现有方法通过架构改进专注于单步分离,采用单输入单输出(SISO)或单输入多输出(SIMO)配置,因为具有混合一致性的音频分离通常被视为一个具有唯一解的回归问题。通过将这些架构扩展为MIMO配置,我们引入了迭代预测,同时不损害架构优势或混合一致性的特性。我们进行了全面的消融研究,将该框架与判别器结合并扩展为生成模型。实验结果表明,将所提出的框架应用于最先进的分离模型时,性能有显著提升。
英文摘要
This paper proposes a general framework for stable and effective iterative audio separation with mixture consistency by extending source separation models to a multi-input multi-output (MIMO) configuration. In the field of audio separation, mixture consistency is an essential property for many applications that require accurate phase and timbral information of target sources. While iterative approaches such as diffusion models achieve perceptually superior results in speech enhancement or user-guided target source separation tasks, most existing methods focus on single-step separation with a single-input single-output (SISO) or single-input multi-output (SIMO) configuration through architectural improvements, since mixture-consistent audio separation is generally regarded as a regression problem that admits a unique solution. By extending these architectures to a MIMO configuration, we introduce iterative prediction without compromising the architectural advantages or the characteristics of mixture consistency. We conduct a comprehensive ablation study of combining the framework with discriminators and extending it to a generative model. Experimental results demonstrate significant performance improvements when applying the proposed framework to state-of-the-art separation models.