发表机构
Tsinghua Shenzhen International Graduate School, Tsinghua University; Tencent(清华大学深圳国际研究生院; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出UniAdapt,通过学习RVQ细化的边际效用并在比特预算下分配音频令牌,实现可扩展的音频表示,显著降低失真并提升语音质量。
AI 中文摘要
离散音频令牌被广泛用作表示接口,然而固定深度的RVQ分词器在每帧上分配相等的容量,尽管各帧的细化价值不同。我们引入了UniAdapt,该方法在冻结的编解码器上学习RVQ细化的边际效用,并在精确的串行化比特预算下分配这些细化。一个与速率无关的因果控制器预测声学效用,而一个可选的语义头支持仅语音的语句级分配;在语音上测得的声学和语义边际增益相关性为0.42。对于因果分配,一个原始-对偶分配器选择前缀有效的深度,而一个精确的守卫将每个序列前缀约束到其匹配的固定深度串行化预算。在语句级分配下,UniAdapt在语音、音乐和环境音频上将Log-STFT失真降低了1.07-4.39个百分点,而无需更大的预算。在因果方式下,它在800次语句-速率评估中改善了四种语音速率中的三种,且零违规,运行速度快于实时。一项20名听者的语句级MUSHRA研究显示,语音有显著的3.52分改善,而在音乐或环境音频上无显著差异。这些结果支持将效用预测与预算执行分离,以实现可扩展的、受预算调节的音频表示。
英文摘要
Discrete audio tokens are widely used as a representation interface, yet fixed-depth RVQ tokenizers allocate equal capacity to every frame despite varying refinement value. We introduce UniAdapt, which learns marginal utility of RVQ refinements on a frozen codec and allocates them under exact serialized-bit budgets. A rate-independent causal controller predicts acoustic utility, while an optional semantic head supports speech-only utterance-level allocation; measured acoustic and semantic marginal gains on speech have a correlation of 0.42. For causal allocation, a primal-dual allocator selects prefix-valid depths, while an exact guard constrains each sequence prefix to its matched fixed-depth serialized budget. Under utterance-level allocation, UniAdapt reduces Log-STFT distortion by 1.07-4.39 percent across speech, music, and environmental audio without larger budgets. Causally, it improves three of four speech rates with zero violations across 800 utterance-rate evaluations and runs faster than real time. A 20-listener utterance-level MUSHRA study shows a significant 3.52-point speech improvement, with no significant differences on music or environmental audio. These results support separating utility prediction from budget enforcement for scalable, budget-conditioned audio representations.
Comments32 pages