发表机构
Dolby Laboratories; University of Maryland, College Park(杜比实验室; 马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VIBE是一种新型文本与视频生成音乐模型,通过深度跨层条件机制与五阶段奖励建模优化,提升了音乐生成的可控性与指令依从性,在生成保真度上与多数基准模型相当。
AI 中文摘要
当前视频生成音乐(V2M)模型缺乏语义控制,且无法惩罚指令违规,这主要是因为它们依赖于重建目标,以及扩散自回归(Diffusion Autoregressive,DAR)架构中静态跨模态条件带来的表征瓶颈。为解决这一问题,我们提出了VIBE,一种新型的文本与视频生成音乐(T+V2M)模型,该模型利用了:(1)条件连接(Conditioning Connection),一种深度跨层条件机制,可动态连接规划头与扩散精调头;(2)综合奖励建模分类法,通过结构化的五阶段训练课程,针对可验证的硬约束(如速度、调式)和主观软质量(如音乐性、多模态对齐)进行优化。在使用视听对齐、指令遵循、音频质量指标,以及主观人类评估研究进行评估后,我们发现VIBE展现出更强的可控性和指令依从性,同时在生成保真度和多模态对齐方面与多数评估的基准模型表现相当。
英文摘要
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.