FRAMEWORKERS:一种用于AI生成视频制作的动态多智能体框架
FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production
AI总结:
FRAMEWORKERS是一种以任务为中心的动态多智能体视频制作框架,通过Director和Assistant分工优化任务编排,在路由准确性、故障恢复等多方面优于现有方案。
AI中文摘要:
现代视频生成器擅长合成单个片段,但完整的视频制作需要协调一长串相互依赖的创意步骤,包括脚本编写、故事板制作、生成和编辑。随着中间输出、依赖关系和执行状态随时间变化,还需要持续的资产管理和动态任务编排。现有的自动化系统通常依赖刚性流程,难以适应不同输入和变化的工作流程;而通用大语言模型(LLM)在长程编排和多模态资产路由方面仍不可靠。我们提出FRAMEWORKERS,一种以任务为中心、基于工作空间的开放式视频制作多智能体框架。中央Director将视频创作建模为动态任务管理,不断编辑任务栈以确定下一个要执行的子任务和要调用的子智能体。Assistant作为执行层,将每个选定任务锚定到共享工作空间,检索所需资产和上下文,调用指定的子智能体,并持久保存生成的工件。执行能力通过带有注册描述符的模块化子智能体公开,允许集成新的子智能体而无需重新设计编排工作流程。为提高编排可靠性,我们通过监督微调(SFT)对Director进行微调,随后使用组相对策略优化(GRPO)进行描述符条件下的任务路由。实验表明,FRAMEWORKERS在路由准确性上优于强大的LLM规划器,能从运行时故障中可靠恢复,无需重新训练即可泛化到未见过的子智能体,且比固定流程、单智能体系统和现有多智能体方法实现更高的端到端视频质量和更广泛的任务覆盖范围。
英文摘要:
Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scripting, storyboarding, generation, and editing. It further demands persistent asset management and dynamic task orchestration as intermediate outputs, dependencies, and execution states evolve over time. Existing automated systems typically rely on rigid pipelines that are difficult to adapt to diverse inputs and changing workflows, while general-purpose large language models (LLMs) remain unreliable for long-horizon orchestration and multimodal asset routing. We introduce FRAMEWORKERS, a task-centric and workspace-grounded multi-agent framework for open-ended video production. A central Director formulates video creation as dynamic task management, continuously editing a Task Stack to determine which subtask to execute next and which sub-agent to invoke. An Assistant serves as the execution layer, grounding each selected task in a shared Workspace, retrieving the required assets and context, invoking the assigned sub-agent, and persisting the resulting artifacts. Execution capabilities are exposed through modular sub-agents with registered descriptors, allowing new sub-agents to be integrated without redesigning the orchestration workflow. To improve orchestration reliability, we fine-tune the Director via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for descriptor-conditioned task routing. Experiments show that FRAMEWORKERS outperforms strong LLM planners in routing accuracy, recovers reliably from runtime failures, generalizes to unseen sub-agents without retraining, and achieves higher end-to-end video quality and broader task coverage than fixed pipelines, single-agent systems, and prior multi-agent approaches.