发表机构
The University of Tokyo; Keio University; University of California, Berkeley(东京大学; 庆应义塾大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文分析13款不同架构的神经音频编解码器在多语料库、多噪声条件下的令牌统计特征,明确声学条件与量化器元类别的主导作用,为该类编解码器的语言统计分析提供架构相关规范。
AI 中文摘要
神经音频编解码器(NACs)将语音转换为离散令牌序列,已有研究表明这些序列遵循类语言统计规律。本文分析13款NACs的令牌统计特征,涵盖多码本残差向量量化(RVQ)、单码本VQ及非VQ设计,在三个语料库的纯净、加性白噪声及真实DEMAND噪声条件下评估。从匹配的令牌样本中估计Zipf参数、Heaps参数、一元熵、码本占用率及Jensen-Shannon散度(JSD),并明确拟合有效性保障与家族条件n元组阶数。语料库身份对所有指标的方差解释度较低,而声学条件与量化器元类别以指标依赖的方式占据主导,其中一元熵是与元类别关联最强的指标。在共同一元阶数下计算的纯净到噪声的JSD,在DEMAND噪声下与梅尔倒谱失真的关联最清晰。先前报道的RVQ编解码器的崩溃与爆炸退化特征,分别集中在白噪声与DEMAND噪声下的RVQ单元;爆炸也发生在非VQ编解码器中,而单码本VQ编解码器则在占用率与分布形状上发生变化,无上述两种特征。这些结果为将语言统计分析应用于NAC令牌提供了架构相关的规范。
英文摘要
Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional $n$-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.
CommentsSubmitted to IEEE Transactions on Audio, Speech, and Language Processing (TASLP)