arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单调性引导的语义对齐用于零样本多说话人图像到语音合成

Monotonicity-Guided Semantic Alignment for Zero-shot Multispeaker Image-to-Speech Synthesis

Lijun Wang, Yixian Lu, Shogo Okada

arXiv 2609.38440首次发表:更新:

AI 中文总结

针对图像到语音合成中的对齐难题,提出MGSA框架,利用语义语音单元和软单调先验实现零样本多说话人合成,在Flickr8k-Audio上表现优异,并支持未见说话人。

AI 中文摘要

直接图像到语音(Img2Sp)在将视觉内容映射到有序语音序列时面临对齐挑战,因为图像允许多种口语描述,并且与语音序列缺乏单调对应关系。我们提出了单调性引导的语义对齐(MGSA),据我们所知,这是首个用于零样本多说话人Img2Sp合成的框架。我们使用语义语音单元来提供跨说话人的共享内容目标,并利用参考语音进行说话人条件设置。查询对齐器通过软单调先验将视觉内容学习到的语义记忆映射到语音单元位置,从而产生位置特定的条件状态。采用块状掩码扩散生成器,在这些状态下进行语音单元生成。在Flickr8k-Audio上的实验显示,与单说话人基线相比,字幕性能具有竞争力,而使用LibriTTS-R参考的评估支持对未见说话人的零样本合成。消融实验验证了对齐器和块扩散的有效性。音频样本可在该https URL获取。

英文摘要

Direct image-to-speech (Img2Sp) poses an alignment challenge in mapping visual content to ordered speech sequences, as images permit multiple spoken descriptions and lack monotonic correspondence with speech sequences. We propose Monotonicity-Guided Semantic Alignment (MGSA), to the best of our knowledge, the first framework for zero-shot multispeaker Img2Sp synthesis. We use semantic speech units to provide shared content targets across speakers with reference speech for speaker conditioning. A query aligner maps semantic memory learned from visual content to speech unit positions via a soft monotonic prior, which yields position-specific conditioning states. A blockwise masked diffusion generator is employed for the speech unit generation conditioning on these states. Experiments on Flickr8k-Audio show competitive captioning performance against single-speaker baselines, while evaluation with LibriTTS-R references supports zero-shot synthesis for unseen speakers. Ablations validate the effectiveness of aligner and block diffusion. Audio samples are available at https://alizeded.github.io/mgsa-demo.

Comments5-pages, 1 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑