arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SepRQ:通过无掩码、多尺度源分离的自监督语音混合表示学习

SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

Séverin Baroudi, Hervé Bredin, Ricard Marxer

arXiv 2610.04690首次发表:更新:

发表机构

Univ Toulon; Aix Marseille Univ; CNRS; pyannoteAI; ILLS(土伦大学; 艾克斯-马赛大学; 法国国家科学研究中心; pyannoteAI; 拉瓦尔国际学习实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SepRQ通过无掩码多尺度源分离目标替代掩码预测,在SUPERB基准上实现说话人日志和语音分离最先进性能,仅需85.68M参数,并开源。

AI 中文摘要

自监督学习(SSL)是语音表示学习的标准方法,但主流模型围绕单说话人音频设计,限制了其在多说话人场景中的实用性。我们提出SepRQ,一个开源SSL框架,用基于冻结随机投影码本的伪源分离目标替代掩码预测。通过采用新颖的无掩码、多分辨率方法,SepRQ在SUPERB基准上的说话人日志和语音分离任务中达到最先进性能,在Base和Large规模上均超越WavLM及其他鸡尾酒会派生的SSL模型,同时仅需85.68M推理参数。SepRQ还在需要注册的目标说话人任务(如目标说话人自动语音识别)以及具有挑战性的多域DIHARD 3日志数据集上展现出强劲性能。值得注意的是,我们报告了在三说话人混合(WSJ0-3Mix)上的强大分离能力,而当前SSL文献在此场景下表现困难。尽管鸡尾酒会SSL仍然稀缺且闭源,仅限于C-HuBERT和基于注册的SA-WavLM,我们将SepRQ开源给社区。

英文摘要

Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective over frozen random-projection codebooks. By adopting a novel mask-free, multiresolution approach, SepRQ achieves state-of-the-art performance in Speaker Diarization and Speech Separation on the SUPERB benchmark, surpassing WavLM and other cocktail-party derived SSLs at both Base and Large scales, while requiring only 85.68M inference parameters. SepRQ also demonstrates strong performance across target-speaker tasks requiring enrollment (such as Target-Speaker Automatic Speech Recognition), and on the challenging multi-domain DIHARD 3 diarization dataset. Notably, we report strong separation capabilities on three-speaker mixtures (WSJ0-3Mix), where current SSL literature struggles. While cocktail-party SSLs remain scarce and closed-source, limited to C-HuBERT and the enrollment-based SA-WavLM, we open-source SepRQ to the community.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑