arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MeanVoiceFlow2:均值流与内容编码器联合优化,实现快速一步零样本语音转换

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo

arXiv 2609.40087首次发表:更新:

发表机构

NTT, Inc.(日本电信电话株式会社)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MeanVoiceFlow2,联合优化均值流与高效内容编码器,通过蒸馏和扩散GAN训练,实现零样本语音转换中更高质量和约9倍加速。

AI 中文摘要

基于流匹配的语音转换(VC)方法因其高语音质量和强说话人相似性而受到关注。其中,诸如MeanVoiceFlow等一步模型因其推理效率高而尤为引人注目;然而,它们对计算密集型内容编码器的依赖仍是瓶颈。为此,我们提出MeanVoiceFlow2,一个联合优化基于流的转换模块和计算高效内容编码器的框架。该模型通过使用MeanVoiceFlow的转换蒸馏和真实数据重建进行训练。我们进一步引入扩散GAN训练,结合样本混合和教师引导的条件增强,以提升真实感和解耦性。在零样本语音转换实验中,MeanVoiceFlow2在保持相当说话人相似性的同时,实现了比MeanVoiceFlow更高的感知质量和约9倍的推理加速。音频样本可在该https URL获取。

英文摘要

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.

CommentsAccepted to Interspeech 2026. Project page: https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑