arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REDnet:用于在说话者数量未知且麦克风数量可变的情况下进行语音分离的递归编码器和解码器

REDnet: Recursive Encoder and Decoder for Speech Separation under Unknown Number of Speakers and Variable Number of Microphones

Fulin Wu, Zhong-Qiu Wang

arXiv 2608.24659首次发表:更新:

AI 中文总结

本研究提出REDnet模型,采用递归编码器和解码器结构,解决了说话者数量和麦克风数量均未知的语音分离问题,在多个公开数据集上实现了领先性能。

AI 中文摘要

我们提出了递归编码器和解码器(RED),用于构建单个深度神经网络(DNN)模型,该模型可分离包含未知数量说话者、未知几何结构下可变数量麦克风的多说话者混合语音,这是一项尚未被研究的任务。RED的解码器会递归检测是否仍有活跃说话者,并一次分离一位说话者,其设计可采用端到端方式训练以提升分离性能。RED的编码器会递归编码输入混合语音的每个麦克风通道,逐步整合空间线索。结合两者,该DNN可被训练以分离不仅说话者数量未知、麦克风数量也可变的混合语音,在多个公共数据集上达到了最先进的性能。

英文摘要

We propose $\textit{recursive encoder and decoder}$ (RED) for building a single deep neural network (DNN) model that can separate multi-speaker mixtures containing unknown numbers of speakers and variable numbers of microphones arranged in an unknown geometry, a task that has not been studied yet. The decoder of RED recursively detects whether there are active speakers left and separates one speaker at a time. It is designed to be trained in an end-to-end fashion to improve separation performance. The encoder of RED recursively encodes each microphone channel of the input mixture, sequentially incorporating spatial cues. Combining both, the DNN can be trained to separate mixtures not only with unknown numbers of speakers but also with variable numbers of microphones, achieving state-of-the-art performance on multiple public datasets.

Commentsin submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑