Poly-InstructTTS:基于开放式指令学习野外场景下的高表现力语音合成
Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
浏览论文内容
中文总结 AI 辅助
针对现有TTS模型难以通过自然语言指令控制细粒度表现力的问题,本文提出Poly-InstructTTS模型,构建1000小时指令标注语料库,采用多模态流水线等技术实现高表现力语音合成,相关成果可在项目页面获取。
中文摘要 AI 辅助
尽管近期的文本转语音(TTS)模型已实现较高的自然度,但通过自然语言指令控制细粒度表现力仍是一项挑战。本文提出Poly-InstructTTS,该模型利用野外场景下的视听数据从开放式指令中学习高表现力语音。我们构建了可扩展的多模态流水线,打造了1000小时的指令标注语料库,涵盖1000余种细粒度情感与风格。该框架采用带属性思考标记的无提示GPT,随后接入注入参考音频音色的流匹配模块。我们还提出说话人微调流程,在保留个性的同时将指令控制迁移至特定说话人。此外,我们扩展了InstructTTSEval,新增更多任务。实验表明,Poly-InstructTTS在指令贴合度与表现力方面表现优异,音频演示与扩展测试集可在项目页面获取。
英文摘要
While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open-ended instructions using in-the-wild audiovisual data. We build a scalable multi-modal pipeline to construct a 1,000-hour instruction-annotated corpus covering 1,000+ fine-grained emotions and styles. The framework uses a prompt-free GPT with attribute-based thinking tokens, followed by a flow-matching module that injects timbre from a reference audio. We also present a speaker fine-tuning procedure to transfer instruction control to specific speakers while preserving persona. We further extend InstructTTSEval with broader tasks. Experiments show that Poly-InstructTTS delivers strong performance in instruction adherence and expressiveness. Audio demos and the expanded testset are available on our project page.
发表机构
- ZuoYeBang Technology(作业帮科技)
机构由 AI 辅助整理,请以论文原文为准。