arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniVAE:用于联合生成的具有跨模态对齐的音频-视频变分自编码器

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu

arXiv 2607.23855首次发表:更新:

AI 中文总结

研究针对音频和视频联合生成中跨模态对齐难的问题,提出OmniVAE,通过联合训练学习音频和视频潜在表示间的细粒度语义对齐,使用对比目标捕捉对应关系并对齐潜在空间,还提炼特征提升可学习性,实验证明其有效提升生成质量和同步准确性。

AI 中文摘要

近期生成模型正从无声视频或独立音频合成迈向同步音频和视频的联合生成。尽管有进展,但因音频和视频的结构差异,实现具有细粒度跨模态对应关系的联合生成仍具挑战。多数现有方法分别训练音频和视频变分自编码器,导致两个潜在空间缺乏跨模态对齐。我们提出OmniVAE,一个联合训练的音频-视频变分自编码器,能学习音频和视频潜在表示之间的细粒度语义对齐。除了重建,OmniVAE使用段级音频-视频对比目标来捕捉时间-语义对应并对齐两个潜在空间。同时,它将预训练的特定模态语义编码器的特征提炼到每个模态中,提高两个潜在空间的下游可学习性。大量实验表明这两个目标一致地提高了潜在空间的可学习性,在下游文本到音频-视频生成中转化为更高的生成质量和更准确的跨模态同步。这些发现强调了学习统一表示作为全模态建模基础的重要性。

英文摘要

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1

Comments15 pages, 2 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑