arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36287eess.AScs.SD

InstCharVoice:在文本到语音中实现自然语言指令的字符级控制

InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

  • Huya Inc.(虎牙公司)

机构由 AI 辅助整理,请以论文原文为准。

Sihang Nie, Xueru Li, Xiaofen Xing, Deyi Tuo, Cheng-Bin Jin, Jingyuan Xing, Jinxin Ji

AI总结:

针对ITTS系统缺乏字符级细粒度控制的问题,提出InstCharVoice框架,利用Qwen3-Omni构建指令标注,训练自回归模型实现指令驱动的字符级声学控制,实验证明其指令遵循和关键词控制优于现有系统。

AI中文摘要:

基于指令的文本到语音(ITTS)系统能够实现对表现力语音生成的自然语言控制,但通常对单个文本单元的透明度和细粒度控制有限。字符级可控的语音合成系统提供了明确的声学控制,但通常依赖于用户指定的声学属性。为弥合这一差距,我们提出了InstCharVoice,一个统一框架,将自然语言指令嵌入字符级声学控制中。我们首先使用Qwen3-Omni在WordVoice-5A-zh语料库上构建了有依据的指令标注。借助这些监督信息,我们训练了一个自回归模型,在生成相应语音标记之前,识别与指令相关的字符并预测其声学属性。关键词预测和基于依据的损失加权有助于模型聚焦于与指令相关的字符和属性。实验表明,与代表性ITTS系统相比,指令遵循和关键词级声学控制得到改善,同时语音自然度具有竞争力且具备明确的字符级可控性。音频样本可在以下网址获取:此https URL。

英文摘要:

Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.

补充信息

↑