发表机构
FPT Software AI Center; NVIDIA Corporation(FPT软件人工智能中心; NVIDIA公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视频到音频生成问题,提出Flowley架构,结合视觉特征与文本提示,通过渐进软掩码交叉注意力实现视听同步,无额外计算成本,还提出SoundCap字幕,该方法在多个指标及零样本音频质量上达先进水平。
AI 中文摘要
视频到音频(V2A)生成旨在合成与无声视频语义一致且时间同步的逼真音频。尽管有进展,但许多方法仍存在多阶段训练成本高、运行时间长或牺牲细粒度时间线索等问题。为此提出Flowley,一种端到端单阶段训练架构,结合视觉特征和文本提示生成音轨。引入渐进软掩码交叉注意力,直接在注意力机制中嵌入视听同步,无额外计算成本。还指出现有V2A基准缺乏声音描述性字幕,提出SoundCap创建详细的声音感知字幕。Flowley在多个指标上实现了最先进性能,结合SoundCap在零样本设置下音频质量超越现有方法。
英文摘要
Video-to-audio (V2A) generation aims to synthesize realistic audio that is both semantically consistent with and temporally synchronized to a silent video. Despite recent progress, many methods still rely on multi-stage training, resulting in high computational costs and long runtimes, or transform visual input into text to leverage pretrained text-to-audio models, sacrificing fine-grained temporal cues. To overcome these limitations, we propose Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts. Crucially, we introduce Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers. We further observe that existing V2A benchmarks lack sound-oriented descriptive captions, which can potentially degrade the quality of the synthesized audio. To remedy this, we propose SoundCap, a plug-and-play pipeline for creating detailed, sound-aware captions that guide the model. Remarkably, without integrating any pretrained audio-visual alignment modules, Flowley achieves state-of-the-art performance on VGGSound across multiple metrics. Moreover, by incorporating SoundCap, we further exceed the performance of the strongest existing close-sourced methods in terms of audio quality in the zero-shot setting.
CommentsAccepted to ECCV 2026