arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在潜在空间中通过跨模态对齐实现精确的视频到音频生成

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran, Long-Khanh Pham, Paarth Neekhara, Shehzeen Hussain, Van Nguyen

arXiv 2607.06405首次发表:更新:

发表机构

FPT Software AI Center; NVIDIA Corporation(FPT软件人工智能中心; NVIDIA公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视频到音频生成问题,提出Flowley架构,结合视觉特征与文本提示,通过渐进软掩码交叉注意力实现视听同步,无额外计算成本,还提出SoundCap字幕,该方法在多个指标及零样本音频质量上达先进水平。

AI 中文摘要

视频到音频(V2A)生成旨在合成与无声视频语义一致且时间同步的逼真音频。尽管有进展,但许多方法仍存在多阶段训练成本高、运行时间长或牺牲细粒度时间线索等问题。为此提出Flowley,一种端到端单阶段训练架构,结合视觉特征和文本提示生成音轨。引入渐进软掩码交叉注意力,直接在注意力机制中嵌入视听同步,无额外计算成本。还指出现有V2A基准缺乏声音描述性字幕,提出SoundCap创建详细的声音感知字幕。Flowley在多个指标上实现了最先进性能,结合SoundCap在零样本设置下音频质量超越现有方法。

英文摘要

Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress, many methods still rely on multi-stage training, resulting in high computational costs and long runtimes, or transform visual input into text to leverage pretrained text-to-audio models, sacrificing fine-grained temporal cues. To overcome these limitations, we propose Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts. Crucially, we introduce Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers. We further observe that existing V2A benchmarks lack sound-oriented descriptive captions, which can potentially degrade the quality of the synthesized audio. To remedy this, we propose SoundCap, a plug-and-play pipeline for creating detailed, sound-aware captions that guide the model. Remarkably, without integrating any pretrained audio-visual alignment modules, Flowley achieves state-of-the-art performance on VGGSound across multiple metrics. Moreover, by incorporating SoundCap, we further exceed the performance of the strongest existing close-sourced methods in terms of audio quality in the zero-shot setting.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑