MambaVoice:基于混合Mamba-Transformer模型的轻量级视听歌声分离
MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model
浏览论文内容
中文总结 AI 辅助
提出轻量级视听歌声分离模型MambaVoice,采用混合Mamba-Transformer架构融合音频与视觉特征,在Acappella和URSing数据集上以1620万参数达到与大型模型相当的性能。
中文摘要 AI 辅助
从音乐视频中分离出目标歌声仍然具有挑战性,尤其是在存在多位歌手和密集器乐伴奏的情况下。我们提出了MambaVoice,一个轻量级的视听框架,利用混合Mamba-Transformer架构进行目标歌声分离。该模型使用基于注意力的频带分割音频编码器和用于面部运动特征的时空图卷积网络(ST-GCN)联合编码音频和视觉流。这些模态通过乘法门控机制融合,使视觉线索能够选择性地调制音频表示。融合后的特征由结合了Transformer自注意力与选择性状态空间模型(SSMs)的混合主干网络处理,以线性复杂度实现高效的长时程建模。我们在Acappella和URSing数据集上评估了MambaVoice,在包括干扰歌手混合物的挑战性条件下。该模型拥有1620万个参数,表现出相当的性能,在Acappella上达到14.18 dB的SDR,并在URSing上展现出强大的跨数据集性能,与参数数量少得多的较大模型相当。这些发现凸显了混合SSM-注意力架构在可扩展、高效的视听源分离中的有效性,表明它们非常适合作为更大流程中的轻量级组件。我们进行了一项感知研究,进一步支持了我们在客观指标上的改进。我们在线上提供了我们的实现。
英文摘要
Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba--Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM--attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.