RVQ位置感知投机解码用于设备端文本到语音
RVQ Position Aware Speculative Decoding for On Device Text to Speech
浏览论文内容
中文总结 AI 辅助
针对设备端TTS中MultiCodeDecoder的RVQ解码瓶颈,提出位置感知投机解码,以极小参数开销实现分布无损加速,将Qwen3-TTS实时合成调用次数从200降至88,RVQ生成提速2至2.2倍。
中文摘要 AI 辅助
自回归解码(AR)与Transformer模型在单流推理时受内存带宽限制,这是设备端文本到语音(TTS)的典型部署场景。使用Qwen3-TTS进行实时流式处理每秒需要超过200次顺序模型调用,其中主要由内循环MultiCodeDecoder主导,该解码器每80毫秒音频帧发射15个残差向量量化(RVQ)码。我们针对MultiCodeDecoder提出RVQ位置感知投机解码,在增加5×10⁻⁴%参数的情况下,每次模型调用接受2.47个令牌,每轮投机/验证开销为10%至20%,将实时合成从每秒200次顺序模型调用减少到88次。该方案在部署的top-k采样下具有分布无损性,与原始系统的词错误率(WER)一致性与此保证相符。我们在最新的iPhone和Apple Silicon Mac设备上,使用Qwen3-TTS 0.6B实现了RVQ令牌生成2至2.2倍的加速。
英文摘要
Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at $5\times10^{-4}$ percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.
发表机构
- Argmax, Inc.(Argmax公司)
机构由 AI 辅助整理,请以论文原文为准。