arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FreyaTTS技术报告

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis

Ahmet Erdem Pamuk, Ömer Yentür, Ahmet Tunga Bayrak, Yavuz Alp Sencer Öztürk, Mustafa Yavuz

arXiv 2607.09530首次发表:更新:

AI 中文总结

介绍无分词器的土耳其语文本转语音模型Freya-TTS,通过无规则端到端建模、非自回归并行去噪等方法推进框架,在基准测试中表现出色,优于开源系统,适合边缘部署,且已开源模型权重、代码和基准测试。

AI 中文摘要

我们介绍了Freya-TTS,这是一个紧凑的、无分词器的、以土耳其语为先的文本转语音模型,专为高度可靠和高效的对话合成而设计。它是一个1.832亿参数的非自回归条件流匹配扩散Transformer(DiT),在AudioVAE2的冻结连续潜在空间中运行。我们在三个关键维度上推进了该框架:无规则端到端建模、非自回归并行去噪、面向生产的两阶段训练后方法。在Freya-TR-Eval基准测试中,Freya-TTS实现了8.0%的带匹配单词错误率和3.0%的字符错误率,优于更大的开源系统,且适合资源受限的边缘部署。我们还发布了模型权重等。

英文摘要

We introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synthesis. Freya-TTS is a 183.2M-parameter non-autoregressive conditional flow-matching Diffusion Transformer (DiT) that operates in the continuous latent space of the frozen AudioVAE2 (16 kHz encode, 48 kHz decode), allowing the model to focus its capacity on text-to-latent mapping while inheriting high-quality 48 kHz reconstruction. We advance the framework along three key dimensions: (1) rule-free end-to-end modeling from a 92-symbol Turkish character vocabulary without a phonemizer, grapheme-to-phoneme frontend, or discrete speech tokenizer, with digit strings expanded to their spoken form at the text frontend; (2) non-autoregressive parallel denoising, which predicts the entire latent sequence simultaneously over a predicted duration; and (3) a production-oriented two-stage post-training recipe consisting of single-speaker voice locking and short-utterance coverage, improving speaker consistency and robustness on short inputs. On the Freya-TR-Eval benchmark, Freya-TTS achieves a band-matched word error rate (WER) of 8.0% and character error rate (CER) of 3.0%, lower error than both larger open systems in its field, XTTS-v2 and F5-TTS, at 40-55% of their parameter count, together with the highest naturalness (MOS) among the compact systems. The model achieves a real-time factor of 0.11 on a consumer GPU (RTX 4090; ~0.14 mean on an H100) and synthesizes in real time on a laptop CPU, making it well suited for resource-constrained edge deployment. We release the model weights, training and inference code, and evaluation benchmark under the Apache-2.0 license.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑