arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27260eess.AS

JASPER:联合音频与语音预训练编码器表示

JASPER: Joint Audio and Speech Pre-trained Encoder Representations

Geeth George, Ameenudeen P E, Hrishikesh H Pillai, Sriram Ganapathy

首次发表
浏览论文内容

中文总结 AI 辅助

JASPER通过向语音预训练模型添加时频目标,实现频谱-时间表示学习,在语音、音频和音乐任务上优于现有基线,验证了统一建模的有效性。

中文摘要 AI 辅助

自监督学习(SSL)在语音和音频领域主要沿各自独立的路径发展:语音模型强调时域预测,而音频表示学习则聚焦于时频模式。这种分离造成了兼容性差距,限制了跨域泛化能力。在这项工作中,我们引入了JASPER(联合音频与语音预训练编码器表示),一个通过时频目标增强语音预训练模型的框架。具体而言,JASPER在长音频片段上对时间和频谱目标进行掩码预测,从而实现对语音和音频信号的频谱-时间表示学习。所提出的方法在多种语音、音频和音乐任务上持续优于多个基线和现有的语音/音频编码器,证明了统一频谱-时间建模的有效性。

英文摘要

Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.

发表机构

  • LEAP Laboratory, Electrical Engineering, Indian Institute of Science(印度科学学院电气工程学院LEAP实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑