arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过几何感知动态卷积实现阵列不变语音增强

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

Zhenglong Liu, Wangyou Zhang, Chenda Li, Yanmin Qian

arXiv 2607.18658首次发表:更新:

发表机构

Auditory Cognition and Computational Acoustics Lab; Shanghai Jiao Tong University; VUI Labs(听觉认知与计算声学实验室; 上海交通大学; VUI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多通道语音增强受固定阵列配置限制问题,提出几何感知动态卷积框架,利用麦克风坐标转换固定阵列SE模型为阵列不变系统,经实验验证该框架能提升模型在不同阵列拓扑下的性能。

AI 中文摘要

多通道语音增强(SE)系统性能优于单通道方法,但受固定麦克风阵列配置限制,阻碍其在不同阵列几何形状设备上的实际部署。近期阵列无关SE方法虽解决了麦克风数量和排列问题,但未充分利用可用的显式阵列几何先验。本文提出几何感知动态卷积(Geo-DConv)框架,利用麦克风坐标将标准固定阵列SE模型转换为强大的阵列不变系统。在RealMAN多通道语音数据集上实验表明,该架构使两个常用固定阵列模型适应阵列不变设置,在不同阵列拓扑中性能持续提升。

英文摘要

Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.

CommentsComments: 5 pages, 1 figure, 2 tables. Accepted at Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑