arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无对齐文本音频盒(Text-AB):用于语音配音和全双工对话合成

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu

arXiv 2609.03992首次发表:更新:

发表机构

FAIR at Meta(Meta FAIR研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无对齐文本音频盒(Text-AB),基于流匹配扩散Transformer,在48万小时语音预训练后经微调,实现高质量配音与全双工对话合成,性能大幅优于现有系统。

AI 中文摘要

我们提出了无对齐文本音频盒(Text-AB),这是一个用于高质量语音配音和全双工对话合成的统一框架。Text-AB基于采用流匹配目标训练的扩散Transformer构建,与Audiobox系统在三个维度上存在差异。首先,它在潜在扩散框架中运行,使用DAC-VAE特征将48 kHz波形编码为25 Hz的潜在序列,相比之前的EnCodec表示实现了10倍以上的压缩,同时提升了重合成质量。其次,Text-AB是无对齐的:它通过现成的文本编码器处理原始文本,并通过交叉注意力学习文本-语音对齐,无需强制对齐和显式时长预测。第三,我们大幅扩展了模型和数据规模,在48万小时的单语语音上预训练了一个30亿参数的模型,随后在三个下游任务上进行监督微调:跨语言语音配音、全双工对话合成以及情感全双工对话合成。推理时,Text-AB支持最多约1分钟语音的单次生成,以及通过多扩散方案实现的任意长文本生成,还有基于自动化指标提升质量的多阶段重排序策略。在真实世界的配音基准上,Text-AB相比最新的内部配音系统实现了阶跃式改进,在韵律相似度、语音相似度、自然度和可分享性方面均有大幅提升。对于全双工对话合成,它在短对话上接近人类录音,在长对话的类人性和表现力上大幅优于最新的内部模型,同时原生建模了话轮转换、反馈回应和情感动态。对于情感对话合成,相比无条件基线,情感 conditioning 显著提升了情感对齐和情感交互质量。

英文摘要

We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑