arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02623cs.SDcs.CLcs.LG

基于语音印象引导的伪三元组构建的可扩展方向跟随TTS

Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

Kenichi Fujita, Yusuke Ijima

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对方向跟随TTS缺乏训练数据的问题,提出伪三元组构建流程,结合可控印象TTS与LLM生成数据,实验验证其可实现稳定的说话人保留修改,结合真实数据可进一步提升方向对齐度

中文摘要 AI 辅助

配音演员常重读同一剧本并根据表演指示调整表达方式,我们将此场景研究为方向跟随TTS,即系统生成反映给定参考话语相对方向的新话语,同时保留说话人身份和语言内容。核心挑战在于缺乏捕捉此类相对修改的训练数据,为此我们提出可扩展的伪三元组构建流程,生成(参考话语、方向文本、修改后话语)三元组;该流程采用可控制印象的TTS模型生成可控风格变化,并利用大语言模型(LLM)从估计的印象差异中生成自然语言方向。实验结果表明,仅伪三元组即可实现稳定的保留说话人修改,结合伪数据与真实记录数据可进一步提升方向对齐度,同时保持说话人相似度。音频示例可在我们的演示页面查看:this https URL

英文摘要

Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/

发表机构

  • NTT, Inc.(NTT公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑