arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPEAR-Gen:面向统一语音表示的生成感知预训练

SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos, Osama Hanna, Philip C. Woodland, Yiting Lu

arXiv 2609.34147首次发表:更新:

发表机构

Meta Superintelligence Labs(元超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SPEAR-Gen,通过任务对齐特征聚合和由粗到细目标实现单一语音表示同时支持理解与生成,在SUPERB和语音重构上验证了其性能提升。

AI 中文摘要

语音理解与生成对语音表示提出了不同的要求,现有模型通常针对其中一种能力进行优化。为缩小这一差距,我们提出了SPEAR-Gen,一种学习单一表示以同时支持两种能力的语音表示模型。任务对齐的特征聚合将冻结编码器中的互补语言学和副语言学信息整合为用于掩码预测的离散目标,而由粗到细的目标结合了对数梅尔重构与残差流匹配,以保留频谱结构和细粒度的声学变化。在SUPERB和语音重构上的实验表明,SPEAR-Gen在保持强大理解性能的同时,显著提升了重构质量和说话人保持度。这些结果证明,单一的语音表示能够有效支持理解与生成两种能力。

英文摘要

Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.

CommentsIn Submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑