arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21128cs.SD

TF-MossFormer:集成卷积门控局部-全局注意力以增强时频域单声道语音分离

TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Shengkui Zhao, Zexu Pan, Haoxu Wang, Biao Tian, Bin Ma, Xiangang Li

AI总结:

研究旨在改进单声道语音分离,提出TF-MossFormer,结合局部与全局注意力,利用内容感知滑动窗口注意力机制及卷积门控,在WSJ0-2Mix数据集上,凭借不同参数设置取得优异的SI-SDRi,性能超越先前方法。

AI中文摘要:

具有全局注意力的Transformer能够捕捉长距离依赖,但可能会忽略对语音分离至关重要的细粒度局部连续性。我们提出了TF-MossFormer,一种时频Transformer,它结合局部和全局注意力,为单声道语音分离联合建模短程和长程上下文。其核心是一个内容感知滑动窗口注意力机制,可动态调整感受野以实现更强的局部交互,避免静态卷积的僵化。与基于时域块的方法不同,TF-MossFormer利用二维频谱图在时间和频率轴上对结构进行建模。注意力层之间的卷积门控进一步改善了特征选择和信息流。TF-MossFormer在WSJ0-2Mix上分别以590万、1690万和2540万个参数实现了22.6、24.0和24.4dB的SI-SDRi,优于先前的方法。

英文摘要:

Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range contexts for monaural speech separation. At its core is a content-aware sliding-window attention mechanism that dynamically adapts receptive fields for stronger local interactions, avoiding the rigidity of static convolutions. Unlike time-domain chunk-based methods, TF-MossFormer leverages the 2D spectrogram to model structure along both time and frequency axes. Convolutional gating between attention layers further improves feature selection and information flow. TF-MossFormer achieves SI-SDRi of 22.6, 24.0, and 24.4 dB on WSJ0-2Mix with 5.9M, 16.9M, and 25.4M parameters, respectively, outperforming prior approaches.

补充信息

↑