基于公共领域电影集的对话感知视频到音乐生成
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
- University of California San Diego(加利福尼亚大学圣地亚哥分校)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对视频到音乐生成的可复现性问题,构建了OSSL-v2数据集,提出用对话作为条件信号的方法,在电影上评估后性能优于现有基线。
AI中文摘要:
视频到音乐生成因在传达包括电影在内的视觉媒体情感方面的作用而受到越来越多的关注。然而,该领域的进展受到可复现性差距的阻碍:模型通常在通过YouTube URL引用的爬取语料库上训练,这些URL可能会被删除,而底层数据往往难以且耗时地检索。为解决这一问题,我们引入了开源银幕原声带库第2版(Open Screen Soundtrack Library version 2,简称OSSL-v2),这是一个自托管的语料库,包含来自公共领域电影的34343个视频片段,总时长为246.4小时。与爬取的语料库不同,OSSL-v2具有可复现性(即不受链接失效影响)且注重版权,同时规模足够大,可用于训练功能正常的视频到音乐模型。然后,我们利用这个电影领域语料库,以电影音乐与屏幕上语音之间的紧密时间耦合为动机,研究对话作为视频到音乐生成的条件信号。具体而言,我们为现有模型的视频交叉注意力增加了一个时间轴,并逐帧用对话轨道对其进行调制。在公共领域电影和商业电影上进行评估时,我们的方法显示出优于最先进基线的性能。该数据集可在此https URL获取。
英文摘要:
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reproducibility gap: models are often trained on crawled corpora referenced through YouTube URLs that may be deleted, with the underlying data often difficult and time-consuming to retrieve. To address this, we introduce the Open Screen Soundtrack Library version 2 (OSSL-v2), a self-hosted corpus of 34,343 video clips totaling 246.4 hours, sourced from public-domain films. Unlike crawled corpora, OSSL-v2 is reproducible (i.e., not subject to link rot) and copyright-conscious, yet still large enough to train functional video-to-music models. We then use this film-domain corpus to study dialogue as a conditioning signal for video-to-music generation, motivated by the close temporal coupling between film music and on-screen speech. Specifically, we augment existing models' video cross-attention with a time axis and modulate it frame-by-frame with the dialogue track. Evaluated on both public-domain and commercial films, our approach shows improvement over the state-of-the-art baselines. The dataset is available at https://huggingface.co/datasets/McAuley-Lab/OSSL-v2.