arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15184cs.SD

跨语言 F5-TTS 2:用于语言无关语音克隆的简化框架

Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

Qingyu Liu, Rixi Xu, Yushen Chen, Zhikang Niu, Haitao Li, Pengcheng Zhu, Bowen Zhang, Jian Zhao, Yunting Yang, Qinyuan Cheng, Xipeng Qiu, Berrak Sisman, Kai Yu, Xie Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出跨语言F5-TTS 2,一种无需强制对齐和转录文本的跨语言语音克隆简化框架,通过预训练模型构造配对数据并微调,增强语速预测器鲁棒性,实现更高说话人相似度。

中文摘要 AI 辅助

零样本文本到语音(TTS)可以从短音频提示中克隆说话人的声音,然而大多数TTS系统在推理时仍需要音频提示的转录文本。这一依赖使得当音频提示的转录文本不可用时,尤其是在未见语言上,无法进行跨语言语音克隆。跨语言F5-TTS消除了这一依赖,实现了无需转录文本的跨语言语音克隆,但它使用强制对齐来准备训练数据。强制对齐对边界错误敏感,并且随着覆盖更多语言,其成本增加。其语速预测器在音频提示以静音开始或结束时,在估计时长方面也不可靠。在本文中,我们提出了跨语言F5-TTS 2,一个无需强制对齐的、用于无转录文本跨语言语音克隆的简化框架。我们不使用强制对齐来分割真实话语,而是利用预训练的F5-TTS模型构建同说话人的提示和目标配对,并在这些构造的配对数据上对同一模型进行微调。这简化了数据准备,并保留了预训练模型的声学建模能力,使得仅通过短暂的微调阶段即可实现适应。我们进一步通过静音感知增强,使音节级语速预测器对前导和尾随静音具有鲁棒性。实验表明,跨语言F5-TTS 2在保持可懂度的同时,达到了比F5-TTS和跨语言F5-TTS更高的说话人相似度。所有相关资源均已公开。

英文摘要

Zero-shot text-to-speech (TTS) can clone a speaker's voice from a short audio prompt, yet most TTS systems still require the audio prompt transcript during inference. This dependency prevents cross-lingual voice cloning when the audio prompt transcript is unavailable, particularly for unseen languages. Cross-Lingual F5-TTS removes this dependency and enables transcript-free cross-lingual voice cloning, but it prepares its training data with forced alignment. Forced alignment is sensitive to boundary errors, and its cost grows as more languages are covered. Its speaking rate predictor is also unreliable at estimating duration when the audio prompt begins or ends with silence. In this paper, we present Cross-Lingual F5-TTS 2, a simplified framework for transcript-free cross-lingual voice cloning without forced alignment. Instead of using forced alignment to segment real utterances, we build same-speaker prompt and target pairs using a pretrained F5-TTS model and fine-tune the same model on these constructed pairs. This simplifies data preparation and preserves the acoustic modeling capability of the pretrained model, enabling adaptation with only a short fine-tuning stage. We further make the syllable-level speaking rate predictor robust to leading and trailing silence through silence-aware augmentation. Experiments show that Cross-Lingual F5-TTS 2 reaches higher speaker similarity than F5-TTS and Cross-Lingual F5-TTS while maintaining intelligibility. All related resources are publicly available.

发表机构

  • Johns Hopkins University(约翰霍普金斯大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创智学院)
  • Geely(吉利)
  • Zhejiang University(浙江大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑