发表机构
National Taiwan Normal University; Academia Sinica(台湾师范大学; 中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对轻量级音视频语音增强中效率与对齐精度的权衡,提出稀疏图引导的Mamba框架,以线性复杂度建模跨模态关系,在LRS3上达13.091 dB SI-SDR,兼顾鲁棒性与低计算成本。
AI 中文摘要
轻量级音视频语音增强(AVSE)模型在计算效率与跨模态对齐准确性之间面临关键权衡。简单的拼接缺乏关系表达能力,而密集的交叉注意力则带来计算开销,并且在强声学干扰下容易产生不可靠的跨模态对应。我们提出稀疏图引导的Mamba(SG-Mamba),这是一种轻量级AVSE框架,将稀疏异构图与线性复杂度的Mamba骨干网络相结合。该图通过内容自适应注意力和跨帧音视频连接显式建模模态特定关系,而Mamba则捕获长时程时间上下文。我们进一步引入音频跳跃连接,以在保留频谱细节的同时不牺牲噪声抑制性能。在LRS3上的评估表明,SG-Mamba在强轻量级基线中取得具有竞争力或更优的性能,并在仅噪声条件下达到13.091 dB的SI-SDR。它在杂乱的多说话人条件下也保持鲁棒性,计算成本为3.45 G MACs(或6.90 G FLOPs),具有竞争力。在VoxCeleb2上的结果进一步表明,显式结构先验提高了轻量级AVSE的鲁棒性、泛化能力和计算效率。
英文摘要
Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.
CommentsAccepted to IEEE SLT 2026