CineForge:用于长时序视频生成的自改进智能体
CineForge: Self-Improving Agents for Long-Horizon Video Generation
- Tongji University(同济大学)
- Kuaishou Technology(快手科技)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出CineForge框架,结合CineForge-Produce与CineForge-Evolve,通过CPPE策略演进机制提升长时序故事驱动视频生成效果,在CineScope指标及基准测试中表现优于现有方法,减少了审查LLM调用量。
AI中文摘要:
长时序故事驱动的视频生成需要一个生产智能体在相互依赖的场景间协调叙事分解、状态跟踪、镜头设计、提示构建、渲染和修订。现有自适应视频系统主要优化请求或可复用技能,而将反复出现的生产故障与跨故事的持续性、阶段针对性改进脱节。我们提出CineForge,一个自演进的视频生产智能体框架,它将用于视频生成的CineForge-Produce与用于跨故事策略演进的CineForge-Evolve相结合。CineForge-Produce将每个源故事组织为类型化叙事、角色、空间和电影状态,利用这些状态协调资产和片段生成,并将过程记录为规范的生产轨迹。CineForge-Evolve应用案例到模式到策略演进(CPPE)来审查轨迹证据,将反复出现的发现整合为有界的阶段局部补丁,并通过结构重放和置信度控制的配对评估部署已验证的更新。为衡量完整故事实现,我们引入CineScope,它结合了包含100个脚本的CineScope-Data套件与人类对齐的多尺度CineScope-Metric,覆盖因果状态、导演编排、节奏与资源分配以及角色弧。在CineScope-Data和两个公共基准上,演进后的CineForge策略将CineScope-Metric从4.024提升至4.380,在ScriptAgent下优于三个长视频基线且增益稳定,并在新故事上减少37.0%的审查LLM调用。这些结果确立生产轨迹为视频智能体的可操作经验,可在长形式叙事任务中持续改进。
英文摘要:
Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems primarily refine requests or reusable skills, leaving recurring production failures disconnected from persistent, stage-targeted improvements across stories. We introduce CineForge, a self-evolving video-production agent framework that couples CineForge-Produce for video generation with CineForge-Evolve for cross-story policy evolution. CineForge-Produce organizes each source story into typed narrative, character, spatial, and cinematic states, uses them to coordinate asset and clip generation, and records the process as a canonical production trajectory. CineForge-Evolve applies Case-to-Pattern-to-Policy Evolution (CPPE) to review trajectory evidence, consolidate recurrent findings into bounded stage-local patches, and deploy validated updates through structural replay and confidence-controlled paired evaluation. To measure complete story realization, we introduce CineScope, which combines a 100-script CineScope-Data suite with a human-aligned, multiscale CineScope-Metric spanning causal state, directorial orchestration, pacing and resource allocation, and character arc. Across CineScope-Data and two public benchmarks, the evolved CineForge policy improves CineScope-Metric from 4.024 to 4.380, outperforms three long-video baselines with consistent gains under ScriptAgent, and reduces review LLM calls by 37.0% on new stories. These results establish production trajectories as actionable experience for video agents that improve cumulatively across long-form storytelling tasks.