arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考语音编解码器:从压缩到自回归生成建模

Rethinking Speech Codecs: From Compression to Autoregressive Generative Modeling

Yazheng Yang, Yao Qiu, Hui Su, Qi Liu

arXiv 2609.04237首次发表:更新:

发表机构

The University of Hong Kong; Meituan Inc.(香港大学; 美团公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将语音令牌化与自回归训练对齐的新型语音编解码器框架,通过引入兼容约束与异构下采样策略,提升了语言模型在语音数据上的持续预训练效果,验证了方法的通用性。

AI 中文摘要

近期语音语言模型的进展利用预训练编解码器得到的离散语音表示,实现可扩展的训练与生成。但现有编解码器主要针对压缩优化,未考虑语言模型训练的自回归特性,导致对压缩语音令牌建模时性能欠佳。本研究从生成建模视角重新审视语音离散化,提出新型框架,明确将语音令牌化与自回归训练对齐。该方法在编解码器训练中引入自回归兼容约束,促使令牌序列具备时间一致性与可预测性;还针对不同层语音令牌提出异构下采样策略,区分语义层与声学层,以提升语义令牌与对应文本内容的对齐度。在多个基准上开展的大量实验表明,本方法缩小了语音压缩与生成建模间的差距,使现有语言模型能在语音数据上开展更有效的持续预训练,且在多种编解码器上均实现性能提升,验证了其通用性及对各类语音建模场景的适用性。

英文摘要

Recent advances in speech language models leverage discrete speech representations from pretrained codecs to enable scalable training and generation. However, existing codecs are primarily optimized for compression without accounting for the autoregressive nature of language model training, resulting in suboptimal performance when modeling compressed speech tokens. In this work, we revisit speech discretization from a generative modeling perspective and propose a novel framework that explicitly aligns speech tokenization with autoregressive training. Our approach introduces autoregressive-compatible constraints during codec training, encouraging token sequences that exhibit temporal consistency and predictability. In addition, we propose a heterogeneous downsampling strategy for different layers of speech tokens, distinguishing semantic from acoustic layers, to improve the alignment between semantic tokens and corresponding textual content. Extensive experiments across multiple benchmarks demonstrate that our method bridges the gap between speech compression and generative modeling, enabling more effective continued pretraining of existing language models on speech data. The approach consistently improves performance across multiple codecs, validating its generality and applicability to diverse speech modeling scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑