发表机构
University of Exeter; Alibaba Group(埃克塞特大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MVAgent通过多智能体协作构建一致条件,并采用镜头级策略优化(Trunk-GDPO),解决多镜头视频生成中的外观、布局和状态一致性问题,在ViMax-Bench上取得最优跨镜头一致性。
AI 中文摘要
多镜头智能体视频生成需要角色外观一致、跨摄像机角度的空间布局稳定以及镜头间角色状态的连续。当每个镜头都是对冻结生成器的单独请求时,重复的文本并不能决定外观、布局或状态。因此,我们将该问题重新定义为条件构建,并提出MVAgent,一个多智能体流水线,其智能体通过类型化条件输入进行协作。由于环境图像仅展示一个视角,空间锚定智能体从生成的摄像机遍历片段中采样视图,并将每个镜头锚定到与其取景匹配的视图上。当生成的镜头偏离计划时,观察者将每个镜头的结束状态记录到连续性记忆中,转换智能体据此为下一个镜头构建角色动作和空间参考。编排者将这些输入组合到每个请求中。由于请求的效果仅在渲染后显现,我们使用Trunk-GDPO的智能体强化学习对其进行训练,该方法在每个镜头处比较渲染候选,而非每视频一次,并继续最佳候选作为主干。在生成器和评判器冻结的情况下,MVAgent在ViMax-Bench上相比所比较的方法获得了最高的跨镜头一致性和叙事规划质量,并在人工评估中优于最强的智能体基线。
英文摘要
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoint, a Spatial Grounding agent samples views from generated camera-traversal clips and anchors each shot to the view matching its framing. As generated shots drift from the plan, an Observer records how each shot ends in a continuity memory, from which a Transition agent builds character action and spatial references for the next shot. An Orchestrator composes these inputs into each request. Since a request reveals its effect only after rendering, we train it by agentic reinforcement learning with Trunk-GDPO, which compares rendered candidates at every shot rather than once per video and continues the best as the trunk. With generator and judges frozen, MVAgent attains the highest cross-shot consistency and narrative-planning quality among the compared methods on ViMax-Bench and is preferred over the strongest agentic baseline in human evaluation.
Comments5 pages, 2 figures