发表机构
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; KlingAI Research; Shanghai Theatre Academy; Beijing Film Academy; University of Konstanz(中国科学院自动化研究所; 中国科学院大学人工智能学院; 可灵AI研究院; 上海戏剧学院; 北京电影学院; 康斯坦茨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MVWeaver提出一种分层音乐视频生成智能体,通过学习歌曲-视觉桥接将歌曲理解转化为镜头计划,利用1,861对真实数据微调LLM,显著提升视觉连贯性与歌曲契合度。
AI 中文摘要
音乐视频是当代文化中一种重要的视听表达形式。它们通过精心的视觉设计来转化并扩展歌曲的表达内容。现有的自动音乐视频(MV)生成系统能够生成视觉上合理的镜头,但往往难以实现长时连贯性和基于歌曲的视觉发展。我们提出了MVWeaver,一种音乐视频生成智能体,它将分层规划与学习型歌曲-视觉桥接相结合,该桥接将歌曲理解转化为可执行的镜头计划。MVWeaver架构包括一个全面的歌曲分析模块、一个构建分层计划的视觉规划器,以及渲染计划内容的下游图像和视频生成模型。为了使通用大语言模型具备特定于MV的歌曲-视觉知识,我们从真实MV衍生的监督信号中学习歌曲分析与视觉规划之间的桥接,并整理了1,861对真实世界的歌曲-MV配对,附有结构化的歌曲侧、MV侧以及教师推断的歌曲-视觉原理标注。利用这些标注,我们对大语言模型进行基于LoRA的监督微调(SFT),以预测指导分层视觉规划的歌曲-视觉桥接。我们的实验表明,该方法在基于歌曲的视觉转换、更丰富的视觉发展以及更高的概念和镜头间连贯性方面表现更强,而消融实验支持了学习型桥接条件化的益处。
英文摘要
Music videos are an important form of audiovisual expression in contemporary culture. They translate and extend the expressive content of songs through deliberate visual design. Existing automatic music video (MV) generation systems can generate visually plausible shots, yet often struggle with long-form coherence and song-grounded visual development. We present MVWeaver, a music video generation agent that integrates hierarchical planning with a learned song-to-visual bridge that translates song understanding into executable shot plans. The MVWeaver architecture comprises a comprehensive song analysis module, a visual planner that constructs hierarchical plans, and downstream image and video generation models that render the planned content. To equip a general-purpose LLM with MV-specific song-to-visual knowledge, we learn a bridge between song analysis and visual planning from real-MV-derived supervision and curate 1,861 real-world song--MV pairs with structured song-side, MV-side, and teacher-inferred song-to-visual rationale annotations. Using these annotations, we perform LoRA-based supervised fine-tuning (SFT) of a large language model to predict song-to-visual bridges that guide hierarchical visual planning. Our experiments demonstrate stronger song-grounded visual translation, richer visual development, and greater conceptual and shot-to-shot coherence, while ablations support the benefits of learned bridge conditioning.
Comments5 pages, 2 figures