将音乐带入视野:音乐驱动的360°视频生成
Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation
- National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
【一句话总结】该研究针对音乐可视化问题,提出情感感知的音乐驱动360°视频生成流程,先预估歌曲情感轨迹,再转化为视觉引导,经SEGA框架生成关键帧,最后合成视频,能反映歌曲情感与结构。
AI中文摘要:
【完整翻译】音乐可视化通过将听觉信号转化为视觉形式,为增强听众对音乐的理解和体验提供了一种强大的方式。然而,大多数现有方法要么严重依赖歌词,要么生成类似于传统音乐视频的平面、非沉浸式视频,这限制了它们传达音乐情感动态和提供沉浸式聆听体验的能力。我们提出了Bring Music The Horizon,这是一种用于音乐驱动的360°视频生成的情感感知管道。给定一首输入歌曲,我们的工作首先通过预测每四个小节级别的效价-唤醒值来估计其情感轨迹。然后,这些值使用EmotiCrafter转换为情感感知视觉引导,并且这些引导向量可以由SEGA框架操纵,该框架为关键帧生成提供细粒度语义控制。最后,将图像到视频模型应用于生成的关键帧,以合成时间上连续的360°视频,用于沉浸式音乐可视化。我们的管道生成反映输入歌曲情感进展和时间结构的360°音乐可视化视频。我们使用来自不同流派的歌曲展示了它的能力,并在我们的项目页面(此https URL)上与代表性的音频到视觉生成基线From-Sound-To-Sight进行了定性比较。
英文摘要:
Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience. We propose Bring Music The Horizon, an emotion-aware pipeline for music-driven 360$^\circ$ video generation. Given an input song, our work first estimates its emotional trajectory by predicting valence-arousal values at the level of every four bars. These values are then converted into emotion-aware visual guidance using EmotiCrafter, and these guidance vectors can be manipulated by the SEGA framework, which provides fine-grained semantic control for keyframe generation. Finally, image-to-video models are applied to the generated keyframes to synthesize temporally continuous 360$^\circ$ videos for immersive music visualization. Our pipeline generates 360$^\circ$ music visualization videos that reflect the emotional progression and temporal structure of the input song. We demonstrate its capability using songs from different genres and provide qualitative comparisons with From-Sound-To-Sight, a representative audio-to-visual generation baseline, on our project page at https://etoile-et-toi-mp3.github.io/BMTH_Project_Page/.