arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GLaS-JEPA:无工程化预测目标的高斯正则化语音自监督学习

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Gaspard Botté, Séverin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber, Xavier Alameda-Pineda, Ricard Marxer, Mirco Ravanelli

arXiv 2609.37798首次发表:更新:

发表机构

Concordia University; CNRS; Univ Toulon; Inria; Univ. Grenoble Alpes; Aix Marseille Univ; Grenoble INP(康考迪亚大学; 法国国家科学研究中心; 土伦大学; 法国国家信息与自动化研究所; 格勒诺布尔阿尔卑斯大学; 艾克斯-马赛大学; 格勒诺布尔国立理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GLaS-JEPA通过直接预测掩码位置的连续表示并采用SIGReg正则化,简化语音自监督学习,在LibriSpeech预训练后显著超越非蒸馏基线。

AI 中文摘要

语音自监督学习旨在为下游语音任务学习通用表示。然而,当前方法依赖于复杂且精心设计的预测目标。我们通过GLaS-JEPA挑战了这一必要性,该框架直接在掩码位置预测当前编码器的连续表示,无需对比学习、离散目标或独立的EMA目标编码器。我们使用SIGReg表示空间正则化防止表示坍缩,从而消除了对工程化目标生成机制的需求。在960小时的LibriSpeech上预训练,我们57M参数的模型在冻结编码器SUPERB ASR上实现了6.89%的词错误率(WER),在槽填充上实现了25.87%的字错误率(CER),分别比最佳的非蒸馏子90M基线提高了43.1%和22.0%。这些结果表明,高度竞争的语音表示可以从一个极其简化的训练方案中涌现出来。

英文摘要

Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑