VideoGen-Agent:强化视频生成智能体
VideoGen-Agent: Reinforcing Video Generation Agents
浏览论文内容
中文总结 AI 辅助
本文提出VideoGen-Agent,一种通过多任务强化学习训练的多模态智能体,协调外部工具以增强视频生成,在VABench基准上显著提升性能,并验证了工具学习与升级的通用性。
中文摘要 AI 辅助
近期视频生成模型的进展使得高保真、时间连贯的视频生成成为可能。然而,这些模型往往难以满足需要专业知识、特定身份、物理一致性或有序事件的提示。在本文中,我们提出了VideoGen-Agent,一个通过多任务智能体强化学习训练的多模态智能体,用于在视频生成中调用外部工具。该智能体通过多轮交互协调增强、生成和验证工具,利用提示和中间观察来指导其决策。我们在一个涵盖六项任务的类别平衡数据集上训练共享策略。在教师生成轨迹上进行监督微调建立了工具使用行为,随后通过强化学习进行优化。一个类别感知的混合奖励评估工具调用的有效性、任务合适的工具使用以及生成视频的质量。我们进一步引入了VABench,一个包含600个提示的保留基准,涵盖程序性知识、单实体和多实体身份保持、物理一致性、场景组成以及多镜头时间结构。在VABench上,VideoGen-Agent相比其基础文本到视频生成器提升了19.1分,从56.5提高到75.6。升级生成工具后,无需额外的智能体训练,分数进一步升至86.1。人类评估者在84.3%的比较中更倾向于升级配置,而非最强的独立基线。这些结果支持在视频生成任务中学习工具使用,并表明训练后的智能体能够受益于生成工具的后续进步。
英文摘要
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools. Project page: https://andyca111.github.io/VideoGen_Agent/
发表机构
- Princeton University(普林斯顿大学)
- Stanford University(斯坦福大学)
- UC Davis(加州大学戴维斯分校)
- MMLab, CUHK(香港中文大学多媒体实验室)
- BenchFlow
- GWU(乔治华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。