arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GAANet:用于音视频语音分离的全局引导非对称注意力网络

GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation

Zhiyuan Zhang, Jingyuan Xu, Yiming Tang, Liu Liu, Dan Guo

arXiv 2610.02752首次发表:更新:

发表机构

Hefei University of Technology(合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对音视频语音分离中多尺度特征融合效率低的问题,提出GAANet,通过非对称多尺度融合和全局引导注意力机制,在LRS2和VoxCeleb2上达到最先进性能,且参数和计算量小。

AI 中文摘要

多尺度设计对于高效的音视频语音分离至关重要,然而,如何有效地建模多尺度信息以进行音视频特征融合仍然是一个挑战。我们认为,现有方法能力受限主要源于两点:1)以相同方式处理来自不同模态的特征;2)忽视了全局特征的作用。为解决这些问题,我们提出了一种全局引导的非对称注意力网络(GAANet)。我们的模型引入了两项核心创新:首先,一个非对称多尺度融合框架,允许音频和视频流在其各自最优的时间分辨率下提取和交互特征,从而无需对称的时间下采样;其次,一种全局引导的注意力机制,将每个模态压缩成一个时间维度为1的紧凑全局令牌,该令牌随后提供高层语义线索,以指导跨尺度的模态内和模态间融合。在LRS2和VoxCeleb2上的实验表明,GAANet达到了最先进的性能,在LRS2上达到了16.5 dB的SI-SNRi,在VoxCeleb2上达到了14.0 dB,同时保持了轻量级的计算特性,仅需3.3M参数和19.8G MACs。这些结果凸显了非对称时间建模和全局引导在高效且稳健的多模态融合方面的巨大潜力。源代码可在该https URL公开获取。

英文摘要

Multi-scale design is crucial for efficient audio-visual speech separation, yet effectively modeling multi-scale information for audio-visual feature fusion remains challenging. We argue that the limited capacity of existing approaches primarily arises from: 1) treating features from different modalities in the same manner, and 2) overlooking the role of global features. To address these issues, we propose a Global-guided Asymmetric Attention Network (GAANet). Our model introduces two core innovations: first, an asymmetric multi-scale fusion framework that allows audio and visual streams to extract and interact with features at their respective optimal temporal resolutions, removing the need for symmetric temporal downsampling; second, a global-guided attention mechanism that compresses each modality into a compact global token with a temporal dimension of one, which then provides high-level semantic cues to guide both intra- and inter-modal fusion across scales. Experiments on LRS2 and VoxCeleb2 demonstrate that GAANet achieves state-of-the-art performance, reaching 16.5 dB SI-SNRi on LRS2 and 14.0 dB on VoxCeleb2, while maintaining a lightweight computational profile with only 3.3M parameters and 19.8G MACs. These results highlight the strong potential of asymmetric temporal modeling and global guidance for efficient and robust multimodal fusion. The source code is publicly accessible at https://github.com/redizzy/GAANet

CommentsAccepted at the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026). 6 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑