发表机构
Concordia University; Mila-Quebec AI Institute; Université Laval(康考迪亚大学; Mila-魁北克人工智能研究所; 拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ZipCodec通过WavLM蒸馏、Transformer架构和球形量化,在6.25Hz超低帧率下实现高质量流式语音编码,显著优于现有编解码器。
AI 中文摘要
神经音频编解码器是现代语音生成系统的基本组成部分。尽管最近的编解码器实现了越来越低的比特率,但降低帧率仍然具有挑战性,因为每个令牌必须在保持重建质量的同时保留更多信息。我们提出了ZipCodec,一种流式神经语音编解码器,以6.25赫兹和0.80千比特每秒的速率运行,理论延迟为160毫秒。我们的方法将大规模WavLM蒸馏与重新设计的基于Transformer的架构、标量球形量化器以及延迟感知的流式解码器相结合。实验表明,在相当的比特率下,ZipCodec在重建和下游任务中均显著优于现有的流式编解码器,同时以更低的帧率运行。尽管拥有8.42亿参数,ZipCodec在消费级CPU上实现了实时单流推理。演示样本、代码和检查点可在该https URL获取。
英文摘要
Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU. Demo samples, code and checkpoints are available at https://lucadellalib.github.io/zipcodec-web/.
Comments5 pages, 1 figure