arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11437cs.SDeess.AS

编辑说话者,控制说话方式:面向TTS的全局音色编辑与局部指令控制

Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for TTS

Junchuan Zhao, Chenglin Xu, Wei Zeng, Haoyang Li, Yiwen Guo, Ye Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出EDICT框架,通过结合全局音色编辑与局部指令控制实现TTS语音合成,在TimbreEdit-Bench等基准上验证了其音色编辑效果及多指标的平衡性能。

中文摘要 AI 辅助

基于指令的文本转语音(TTS)通过语音克隆、基于文本的语音设计等接口实现对语音特征与表达的控制。语音克隆可复现参考语音,基于文本的语音设计则能通过自然语言描述生成语音,但这两种接口均无法让用户直接修改给定参考语音的音色并合成修改后的语音;同时,语句级表达指令对单个文本片段的修改规定不够明确。本文提出EDICT框架,其通过编辑后的声学参考在各片段间锚定语音身份,统一全局音色编辑与局部表达控制。为实现指令编辑语音的合成,EDICT将参考音频与结构化音色编辑结合,在编码令牌空间生成编辑后的参考,该表示作为冻结TTS主干的共享语音锚点,允许片段级自然语言指令引导表达。为适配指令变化并支持声学连续性,EDICT在每个片段边界重建KV缓存,更新指令条件的同时保留先前生成语音的有限声学上下文。在本文提出的TimbreEdit-Bench与IntraTTS-Bench上的评估表明,EDICT的音色编辑效果有所提升,且在局部指令依从性、说话人一致性与过渡质量间取得了良好平衡。音频演示已公开。

英文摘要

Instruction-based text-to-speech (TTS) offers control over voice characteristics and speech expression through interfaces including voice cloning and text-based voice design. Voice cloning reproduces a reference voice, whereas text-based voice design creates a voice from a natural-language description. However, neither interface directly enables users to modify the timbre of a given reference and synthesize speech with the modified voice. Meanwhile, utterance-level expressive instructions leave changes across individual text segments underspecified. We introduce \textbf{EDICT}, a framework that unifies global timbre editing and local expressive control by using an edited acoustic reference to anchor voice identity across segments. To enable synthesis with an instruction-edited voice, EDICT combines reference audio with structured timbre edits to generate an edited reference in codec-token space. This representation serves as a shared voice anchor for a frozen TTS backbone, allowing segment-specific natural-language instructions to guide expression. To accommodate instruction changes while supporting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary, refreshing instruction conditioning while retaining bounded acoustic context from previously generated speech. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre editing and a favorable balance between local instruction adherence, speaker consistency, and transition quality. Audio demos are available.

发表机构

  • National University of Singapore(新加坡国立大学)
  • LIGHTSPEED(光速(机构名))
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑