发表机构
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VoiceWeaver通过结构化标签前缀和分阶段训练,在共享文本-音频模型中实现表现力语音和声音事件生成,保留先前能力,提升多属性控制准确率。
AI 中文摘要
为语音生成器添加表现力和环境控制需要学习异构属性,同时不丧失先前的能力。VoiceWeaver通过结构化标签前缀和分阶段情感-音调-事件训练,在共享的文本-音频模型中解决这一问题。基于重放的蒸馏保留了先前的预测,而嵌入去相关和属性丢弃则对条件进行正则化。评估区分了单一属性正确性、联合情感-事件生成以及三属性正确性。四舍五入到整数百分比,中文情感准确率为86%,英文为83%。中文情感-事件联合准确率约为70%,而混合任务训练为56%,Ming-omni-tts为48%;英文联合准确率为66%。三属性准确率在中文和英文各500个样本上分别为59%和55%(合并为57%)。文本错误率超过外部TTS基线。音频样本可在该https URL获取。
英文摘要
Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.