发表机构
Beijing University of Posts and Telecommunications; Li Auto(北京邮电大学; 理想汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Instruct-TTS中指令监督的语义漂移问题,提出可控多样化、漂移过滤与属性对齐三机制数据方案,将指令跟随率提升至56.4%,漂移率降至15.4%。
AI 中文摘要
Instruct-TTS系统通过LLM重写将结构化风格标签扩展为自然语言训练指令,然而我们发现超过40%的无约束重写包含语义漂移,这会破坏监督并削弱泛化能力。我们将此问题形式化为指令监督不稳定性,并提出一种以数据为中心的稳定化方案,通过三种机制联合提升覆盖度与保真度:用于系统性扩展的可控指令多样化、用于质量控制的基于LLM的漂移过滤,以及将韵律控制锚定于声学扰动的属性对齐监督。在InstructTTSEval的中文子集上,我们的方案将指令跟随率从无微调时的34.5%和朴素微调时的51.0%提升至56.4%,同时约束重写将漂移率从40.4%降至15.4%。消融实验证实三种机制具有互补性,且该漂移分类法可能推广至TTS之外的指令驱动生成任务。
英文摘要
Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.
Comments5 pages, 2 figures, 5 tables. Audio demos: https://piedpiperg.github.io/instruct-tts-stabilizer/#audio-demos