发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LACE提出了一种层自适应动态帧率音频编解码器,通过每层独立压缩和联合对齐机制,在重建任务上取得更优的率-质量权衡,并提升TTS推理效率。
AI 中文摘要
神经音频编解码器是语音语言建模中的关键组成部分。然而,其高帧率导致序列长度较长,从而增加了计算成本。动态帧率编解码器通过使用压缩步骤将多个帧合并在一起来降低有效帧率,从而缓解这一问题。然而,大多数先前的方法要么仅适用于单码本编解码器,要么在多层级量化之前应用单个压缩步骤。这迫使所有量化层共享相同的分段边界,尽管不同量化层的残差嵌入随时间表现出不同的变化率。我们提出了LACE(层自适应编解码器编码),这是一种动态帧率编解码器,在每个量化层应用独立的压缩步骤,从而实现特定于层的分段边界。为了在下游文本到语音(TTS)任务中使用LACE令牌,我们进一步引入了联合对齐和边界锚定机制,以使各层之间的持续时间保持一致,同时保留压缩优势。在LibriTTS上的实验表明,LACE在重建任务上比先前的动态帧率方法提供了更好的率-质量权衡,并提高了TTS推理效率,同时保持了有竞争力的合成质量。我们的代码作为ESPnet3编解码器配方的一部分发布。
英文摘要
Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.
CommentsAccepted to SLT 2026. 8 pages, 5 figures