发表机构
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UniStream提出多专家残差向量量化(ME-RVQ)和OT-CFM辅助目标,实现48 kHz因果流式音频编码,在12/22.5 kbps下超越现有系统。
AI 中文摘要
我们提出了UniStream,一种用于流式语音、音乐和环境声音的全因果48 kHz神经音频编解码器。其核心是多专家残差向量量化(ME-RVQ),该机制将每个残差量化层中的单一共享码本替换为由确定性Top-K路由器控制的四个专家码本。由于路由决策仅基于先前解码的量化状态得出,解码器无需传输专家标识符即可重现所选专家,从而在增加5.5M参数的同时扩展了量化容量。我们进一步引入辅助的最优传输条件流匹配(OT-CFM)目标,以在训练期间正则化量化潜空间。流模块在推理时被完全移除,因此不产生运行时开销。UniStream在因果48 kHz编码器-解码器框架内支持12 kbps的Top-1模式和22.5 kbps的Top-2模式,同时实现实时GPU推理。为补充窄带语音指标,我们报告了48 kHz ViSQOL音频模式、ViSQOL语音模式、标准VGGish-FAD、DNSMOS P.835以及与Opus和EnCodec的更高比特率参考比较。在12 kbps下,UniStream-Top1的PESQ和UTMOS得分与EnCodec相当,同时将语音Mel-D从13.07降至8.21。在22.5 kbps下,UniStream-Top2的ViSQOL语音模式得分为4.67,环境音频模式得分为3.96,在后一设置中超过了所有在12 kbps或以下运行的系统。在ViSQOL音频模式下,其语音MOS-LQO得分与24 kbps Opus相差0.03以内。消融研究证实,ME-RVQ是质量提升的主要来源,而OT-CFM在语音上提供了感知增益,但以轻微的频谱失真为代价。
英文摘要
We present UniStream, a fully causal 48 kHz neural audio codec for streaming speech, music, and environmental sounds. At its core is Multi-Expert Residual Vector Quantization (ME-RVQ), which replaces the single shared codebook in each residual quantization layer with four expert codebooks controlled by a deterministic Top-K router. Because routing decisions are derived solely from previously decoded quantized states, the decoder can reproduce the selected experts without transmitting expert identifiers, thereby expanding quantization capacity while adding 5.5M parameters. We further introduce an auxiliary Optimal Transport Conditional Flow Matching (OT-CFM) objective to regularize the quantized latent space during training. The flow module is removed entirely at inference and therefore incurs no runtime overhead. UniStream supports a 12 kbps Top-1 mode and a 22.5 kbps Top-2 mode within a causal 48 kHz encoder-decoder framework, while achieving real-time GPU inference. To complement narrow-band speech metrics, we report 48 kHz ViSQOL audio mode, ViSQOL speech mode, standard VGGish-FAD, DNSMOS P.835, and higher-rate reference comparisons with Opus and EnCodec. At 12 kbps, UniStream-Top1 achieves PESQ and UTMOS scores comparable to EnCodec while reducing speech Mel-D from 13.07 to 8.21. At 22.5 kbps, UniStream-Top2 achieves a ViSQOL speech-mode score of 4.67 and an environmental audio-mode score of 3.96, exceeding all evaluated systems operating at 12 kbps or below in the latter setting. It also comes within 0.03 MOS-LQO of Opus at 24 kbps on speech in ViSQOL audio mode. Ablation studies confirm that ME-RVQ is the primary source of quality improvement, whereas OT-CFM provides perceptual gains on speech with a mild trade-off in spectral distortion.
Comments15 pages, 1 figure, 5 tables. Accepted at NCMMSC 2026