arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17900cs.SD

利用TTS:借助利用层实现上下文感知的富有表现力的语音合成

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

Shengfan Shen, Di Wu, Xingchen Song, Dinghao Zhou, Pengyu Cheng, Sixiang Lyu, Jian Luan, Shuai Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对语音助手富有表现力的语音合成中灵活风格控制的需求,提出Harness TTS轻量级控制层,通过封闭集提示 - 工具路由实现风格控制,实验证明其在路由和合成任务中表现出色,为语音助手表达控制提供实用方案。

中文摘要 AI 辅助

语音助手的富有表现力的语音合成需要灵活的风格控制,以适应明确请求和更广泛的交互上下文。我们提出了Harness TTS,这是一个轻量级控制层,围绕TTS引擎以外部化并管理其表达行为。它将风格控制重新制定为封闭集提示 - 工具路由:离线时,使用结构化元数据构建风格提示工具的紧凑注册表;在线时,一个大语言模型规划器根据优先级感知观察模式选择合适的工具,TTS执行器使用相应的提示音频合成语音。我们在路由和合成任务上评估了Harness TTS。在路由方面,Qwen3 - 4B在显式、隐式和冲突子集上的Top - 1准确率分别为74.3%、43.0%和64.6%。对于合成,在CosyVoice3和VoxCPM2上的实验表明,Harness TTS优于仅指令控制,实现了更高的指令跟随胜率(在CosyVoice3上优势为23.1 - 35.6分,在VoxCPM2上为13.8 - 20.0分),并将UTMOSv2分数提高了0.11 - 0.38。此外,4B规划器在标准模式下不到50毫秒就能给出第一个工具推荐,对实时交互引入的延迟可忽略不计。这些结果表明,为TTS引擎配备专用的利用层为语音助手表达控制提供了一个实用、可审计且上下文感知的解决方案。

英文摘要

Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightweight control layer that wraps around a TTS engine to externalize and govern its expressive behavior. It reformulates style control as closed-set prompt-tool routing: offline, a compact registry of stylistic prompt tools is constructed with structured metadata; online, an LLM planner selects the appropriate tool based on a priority-aware observation schema, and the TTS executor synthesizes speech using the corresponding prompt audio. We evaluate Harness TTS on both routing and synthesis tasks. In routing, Qwen3-4B achieves Top-1 accuracies of 74.3%, 43.0%, and 64.6% on explicit, implicit, and conflict subsets. For synthesis, experiments on CosyVoice3 and VoxCPM2 show that Harness TTS outperforms instruction-only control, achieving higher instruction-following win rates (margins of 23.1-35.6 points on CosyVoice3 and 13.8-20.0 points on VoxCPM2) and improving UTMOSv2 scores by 0.11-0.38. Moreover, the 4B planner delivers its first tool recommendation in under 50 ms in standard mode, introducing negligible latency for real-time interaction. These results demonstrate that equipping TTS engines with a dedicated Harness layer offers a practical, auditable, and context-aware solution for voice assistant expression control.

发表机构

  • Xiaomi Inc.(小米公司)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑