发表机构
ByteDance Seed(字节跳动 Seed)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自回归语音生成的序列长度、表征容量与长时稳定性平衡难题,本文提出Locodec编解码器与MP-ELD流匹配框架,实现高保真、高可预测性且稳定的长时语音合成。
AI 中文摘要
平衡序列长度、表征容量与长时稳定性是自回归(AR)语音及音频生成的核心问题。高帧率或高容量的表征可保留更多信号细节,但会使流式生成更易受分布漂移和AR误差累积影响;反之,更短、更压缩的表征简化了AR建模,但其有限带宽可能丢弃重要成分,限制重建保真度与生成质量的上限。本文探究低帧率、高维、高带宽的连续表征能否与流式生成框架协同设计,以支持鲁棒高保真重建、强单标记可预测性及优异长时稳定性。本文将该目标分解为两个耦合问题:高维表征空间应具备何种几何与统计特性,以及应如何构建AR连续标记生成器以抵御误差累积。据此,本文提出Locodec,一种局部编码编解码器,其塑造的表征空间可提升低维核心流形的插值性与原生高维坐标的可识别性,进而提升高维高带宽标记的可预测性;还提出MP-ELD,一种单标记AR流匹配框架,采用多路径信息路由与无分类器残差引导以缓解误差累积。对8Hz、768维标记的实验表明,本文设计在不使用外部SSL/ASR模型、预训练文本语言模型或后训练阶段的情况下,既保留了重建质量,提升了单标记可预测性,达到了有竞争力的词错误率(WER),又实现了稳定的长音频合成。
英文摘要
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with higher frame rates or greater capacity can preserve more signal detail, but they also make streaming generation more vulnerable to distribution drift and AR error accumulation. Conversely, shorter and more compressed representations simplify AR modeling, but their limited bandwidth may discard important components and constrain the upper bound of reconstruction fidelity and generation quality. We ask whether a low-frame-rate, high-dimensional, high-bandwidth continuous representation can be co-designed with a streaming generation framework to support robust high-fidelity reconstruction, strong single-token predictability, and superior long-horizon stability. We decompose this goal into two coupled problems: what geometric and statistical properties a high-dimensional representation space should have, and how an AR continuous-token generator should be structured to resist error accumulation. Accordingly, we propose Locodec, a locally encoded codec that shapes its representation space to improve the interpolatability of a lower-dimensional core manifold and the identifiability of the native high-dimensional coordinates, thereby improving the predictability of high-dimensional high-bandwidth tokens. We also propose MP-ELD, a single-token AR flow-matching framework that uses multi-path information routing and residual classifier-free guidance to mitigate error accumulation. Experiments with 8-Hz, 768-dimensional tokens show that our design preserves reconstruction quality, improves single-token predictability, achieves competitive WER, and maintains stable long-form synthesis, without using external SSL/ASR models, pretrained text language models, or post-training stages.