arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACodec:一种采用并行加性矢量量化的低码率神经语音编解码器

PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

Fei Liu, Yang Ai, Xiao-Hang Jiang, Zhen-Hua Ling

arXiv 2609.03363首次发表:更新:

发表机构

National Engineering Research Center of Speech and Language Information Processing; University of Science and Technology of China(国家语言与语音信息处理工程研究中心; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于并行加性矢量量化(PAVQ)的PACodec,采用“全局-局部-全局”设计,在相同解码质量下比基线降低30%码率,且具备解纠缠友好性,可应用于语音转换等下游任务。

AI 中文摘要

本文提出了PACodec,一种基于并行加性矢量量化(PAVQ)的新型低码率神经语音编解码器。与大多数神经语音编解码器中使用的主流残差矢量量化(RVQ)不同,残差矢量量化中矢量量化器(VQs)是顺序依赖的,PACodec采用的PAVQ策略聚合并行量化结果以优化码率使用。具体而言,PAVQ采用“全局-局部-全局”(GLG)设计:全局编码特征由多个独立的VQs并行量化,每个VQ关注表示的局部分量,其输出通过加法聚合以产生用于解码的最终全局量化结果。实验结果表明,由于每个VQ仅关注局部信息,PACodec支持更小的码本,在相同解码质量下与基线相比降低了30%的码率,且仅带来极小的模型复杂度。进一步分析显示,由于PAVQ的GLG框架,所提出的PACodec具备解纠缠友好性,每个独立的VQ捕捉语音的不同方面,例如内容、音色和声学细节,这表明其在语音转换等下游任务中具有应用潜力。

英文摘要

This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregates parallel quantization results to optimize bitrate usage. Specifically, the PAVQ adopts a "global-local-global" (GLG) design: the global encoded features are quantized in parallel by multiple independent VQs, each attending to a local component of the representation, and their outputs are aggregated through addition to yield the final global quantization result for decoding. Experimental results show that PACodec, as each VQ focuses only on local information, supports smaller codebooks and reduces bitrate by 30% compared with baselines at the same decoding quality, with only minor model complexity. Further analysis shows that, owing to the GLG framework of PAVQ, the proposed PACodec is disentanglement-friendly, and each independent VQ captures different aspects of speech, e.g., content, timbre, and acoustic details, suggesting potential for application to downstream tasks such as voice conversion.

CommentsAccepted by APSIPA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑