arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

建模、缩放与解码:利用非语言发声优化可控语音生成

Modeling, Scaling, and Decoding: Optimizing Controllable Speech Generation with Nonverbal Vocalizations

Ziyu Zhang, Yun Chen, Taihui Wang, Hanzhao Li, Qicong Xie, Rilin Chen, Zhixian Zhao, Lei Xie

arXiv 2609.14231首次发表:更新:

发表机构

Tencent HY Speech Team(腾讯HY语音团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非语言发声可控合成的挑战,提出NVV感知的DiTAR系统,通过专用令牌编码、针对性数据增强与多指标采样优化,在ISCSLP 2026挑战赛中取得总体第一。

AI 中文摘要

非语言发声(NVV)的可控合成对于自然且富有表现力的语音至关重要,但由于其声学多样性及现有语料库中的分布不均衡,这一任务仍具挑战性。为应对这些挑战,我们开发了一个NVV感知的DiTAR系统,该系统对连续语音潜变量进行建模,将16种目标NVV类别编码为专用令牌,并调整停止预测以区分话语中部的发声与话语边界。训练首先在大规模双语NVV语音上进行预训练,随后在通过针对性合成增强和频率感知再平衡增强的语料库上继续进行监督微调。在推理时,我们选择声学提示,调整LM引导和噪声注入的尺度,并应用带多指标选择的Best-of-N采样以减少生成失败。最终系统在ISCSLP 2026 NVVSpeech挑战赛Track 2中取得了62.786的官方加权双语得分,在普通话中排名第一,英语中排名第二,并在所有参赛系统中总体排名第一。消融研究表明,针对性增强对代表性不足的NVV类别收益最大,而稳健的候选选择需要在NVV正确性、词汇保真度和感知质量之间取得平衡。

英文摘要

Controllable synthesis of nonverbal vocalizations (NVVs) is es- sential for natural and expressive speech, but remains challeng- ing due to their acoustic diversity and imbalanced distribution in existing corpora. To address these challenges, we develop an NVV-aware DiTAR system that models continuous speech latents, encodes the 16 target NVV categories as dedicated to- kens, and adapts stop prediction to distinguish mid-utterance vocalizations from utterance boundaries. Training begins with large-scale bilingual pre-training on diverse NVV speech, fol- lowed by continued supervised fine-tuning on a corpus en- hanced through targeted synthetic augmentation and frequency- aware rebalancing. At inference time, we select the acoustic prompt, tune the LM-guidance and noise-injection scales, and apply Best-of-N sampling with multi-metric selection to re- duce generation failures. The final system achieves an official weighted bilingual score of 62.786, ranking first in Mandarin, second in English, and first overall among participating systems in Track 2 of the ISCSLP 2026 NVVSpeech Challenge. Ab- lation studies show that targeted augmentation benefits under- represented NVV categories the most, while robust candidate selection requires balancing NVV correctness, lexical fidelity, and perceptual quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑