arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31810cs.SDcs.AIcs.CVcs.MM

游戏视频到音乐的生成

Video-to-Music Generation for Gameplay Videos

Felipe Marra, Lucas N. Ferreira

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种用于游戏视频的音乐生成模型,利用新数据集和冻结ViViT编码器,在客观指标上超越现有基线,参数减少18%。

中文摘要 AI 辅助

近年来,视频到音乐模型取得了显著进展,尤其是在电影和音乐视频应用中。在本文中,我们研究了该问题在电子游戏领域中的应用,这为这些模型带来了新的挑战:视频帧是渲染图形,音乐主要是合成音频,配乐在整个关卡中循环播放,而不是跟随屏幕上的事件。我们引入了一个新的数据集,包含217.6小时的超级任天堂(SNES)游戏视频,配以485小时的干净配乐(无音效和配音),并通过音频指纹识别与游戏音频匹配。利用该数据集,我们训练了一个简单的编码器-解码器Transformer,将视频特征直接传递给MusicGen解码器,比较了不同的编码策略:文本描述(T5)、独立帧(ViT)或时空补丁(ViViT)。每个编码器在冻结和微调两种情况下进行测试,而解码器始终进行微调。冻结的编码器在所有指标上匹配或优于其微调版本,并且冻结的ViViT取得了最佳总体结果。我们使用客观指标和听力研究(N=96)将该模型与最先进的基线进行比较。尽管参数减少了多达18%,我们的模型在客观指标上优于所有基线,在听力研究中超越了GVMGen,并与OSSL表现相当。

英文摘要

Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of clean soundtracks, free of sound effects and voice-overs, matched to gameplay audio via audio fingerprinting. With this dataset, we train a simple encoder-decoder transformer that passes video features directly to a MusicGen decoder, comparing different encoding strategies: textual descriptions (T5), independent frames (ViT), or spatiotemporal patches (ViViT). Each encoder is tested both frozen and fine-tuned, while the decoder is always fine-tuned. Frozen encoders match or outperform their fine-tuned counterparts on every metric, and the frozen ViViT achieves the best overall results. We compare this model with state-of-the-art baselines using both objective metrics and a listening study (N = 96). Despite having up to 18% fewer parameters, our model outperforms all baselines on objective metrics, surpasses GVMGen in the listening study, and performs comparably to OSSL.

补充信息

↑