CycleSpeech:面向指令控制语音合成与副语言理解的互惠对齐
CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding
浏览论文内容
中文总结 AI 辅助
CycleSpeech通过结构化语音画像连接语音合成与理解,利用前向和后向循环互惠反馈及CycleGRPO策略优化,在双语基准上显著提升指令遵循与画像恢复,无需人工偏好标注。
中文摘要 AI 辅助
指令控制的语音合成与副语言理解通常被独立训练,导致这两个任务之间的互惠反馈未被充分探索。我们引入了CycleSpeech,一个通过共享的、结构化的语音画像连接生成与理解的框架,该画像作为监督和互惠反馈的共同目标。前向循环通过比较恢复的画像与目标画像,评估合成语音是否表达了预期的属性。后向循环评估从真实语音中推断出的画像能否指导源说话风格的 reconstruction(重建)。为支持这两个方向,我们构建了一个包含20,046个样本的双语数据集,将指令、目标语音、说话人参考和结构化画像配对。在联合监督微调的基础上,CycleGRPO通过基于画像一致性和说话风格重建的互惠奖励交替更新策略。固定的目标画像锚定来自不断演变的对应方的反馈。该过程既不需要人工偏好标注,也不需要额外的偏好训练奖励模型。在中文和英文基准上的评估显示,指令遵循和画像恢复得到改善,同时保持了有竞争力的合成质量。与Step-Audio-2-mini相比,CycleSpeech在中文和英文上的指令匹配准确率分别提高了4.50和10.06个百分点。受控消融进一步支持了循环反馈对生成控制的贡献。这些结果支持将结构化语音画像作为语音生成与副语言理解之间互惠训练的接口。在线演示可在该https URL获取。
英文摘要
Instruction-controlled speech synthesis and paralinguistic understanding are often trained independently, leaving reciprocal feedback between the two tasks underexplored. We introduce CycleSpeech, a framework that connects generation and understanding through a shared, structured voice profile that serves as a common target for supervision and reciprocal feedback. The forward cycle assesses whether synthesized speech expresses the intended attributes by comparing recovered and target profiles. The backward cycle evaluates whether profiles inferred from real speech can guide reconstruction of the source speaking style. To support both directions, we construct a bilingual dataset of 20,046 examples pairing instructions, target speech, speaker references, and structured profiles. Building on joint supervised fine-tuning, CycleGRPO alternates policy updates using reciprocal rewards grounded in profile consistency and speaking-style reconstruction. Fixed target profiles anchor feedback from the evolving counterpart. This procedure requires neither human preference annotations nor an additional preference-trained reward model. Evaluations on Chinese and English benchmarks show improved instruction adherence and profile recovery while maintaining competitive synthesis quality. Compared with Step-Audio-2-mini, CycleSpeech improves instruction-match accuracy by 4.50 and 10.06 percentage points in Chinese and English, respectively. Controlled ablations further support the contribution of cycle feedback to generation control. These results support structured voice profiles as an interface for reciprocal training between speech generation and paralinguistic understanding. An online demo is available at https://cyclespeech.github.io.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Tsinghua University(清华大学)
- Amphion Technology Co., Ltd.(Amphion科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。