发表机构
Concordia University; CNRS; Univ Toulon; Inria; Univ. Grenoble Alpes; Aix Marseille Univ; Grenoble INP(康考迪亚大学; 法国国家科学研究中心; 土伦大学; 法国国家信息与自动化研究所; 格勒诺布尔阿尔卑斯大学; 艾克斯-马赛大学; 格勒诺布尔国立理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GLaS-JEPA通过直接预测掩码位置的连续表示并采用SIGReg正则化,简化语音自监督学习,在LibriSpeech预训练后显著超越非蒸馏基线。
AI 中文摘要
语音自监督学习旨在为下游语音任务学习通用表示。然而,当前方法依赖于复杂且精心设计的预测目标。我们通过GLaS-JEPA挑战了这一必要性,该框架直接在掩码位置预测当前编码器的连续表示,无需对比学习、离散目标或独立的EMA目标编码器。我们使用SIGReg表示空间正则化防止表示坍缩,从而消除了对工程化目标生成机制的需求。在960小时的LibriSpeech上预训练,我们57M参数的模型在冻结编码器SUPERB ASR上实现了6.89%的词错误率(WER),在槽填充上实现了25.87%的字错误率(CER),分别比最佳的非蒸馏子90M基线提高了43.1%和22.0%。这些结果表明,高度竞争的语音表示可以从一个极其简化的训练方案中涌现出来。
英文摘要
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.