SPIRAL:通过反思规划代理实现自演化动作条件视频生成
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
- Zhejiang University(浙江大学)
- KnowledgeXLab at Shanghai AI Lab(上海人工智能实验室知识X实验室)
- National University of Singapore(新加坡国立大学)
- Chinese Academy of Sciences(中国科学院)
- Tencent Youtu Lab(腾讯优设实验室)
- Nanyang Technological University(南洋理工大学)
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出SPIRAL框架,通过反思规划代理实现长时域动作条件视频生成,解决传统方法在长时间视频生成中的不足,通过闭环设计和自演化机制提升视频生成的一致性和准确性。
中文摘要 AI 辅助
长时域动作条件视频生成旨在合成符合复杂动作指令的时序一致视频,要求过程有序、持续执行动作和场景一致,超越传统TI2V的短时精度。现有单次视频生成模型通常采用开环方式,导致动作执行不完整、幻觉运动和时间漂移。为解决此问题,我们提出SPIRAL,一种闭环框架,通过顺序规划和迭代反思进行动作条件长时域视频生成。具体而言,SPIRAL实现一个思考-行动-反思过程:PlanAgent将高层目标分解为子动作,这些动作条件VideoGenerator生成每个片段并伴随记忆上下文,同时CriticAgent评估中间视频片段以提供迭代优化的反馈。此闭环设计进一步通过利用PlanAgent提出的行为和CriticAgent得出的奖励进行GRPO基于的后训练,以增强视频生成器的长时域一致性。此外,我们引入ActVideoGen-Dataset用于任务特定训练,并建立ActVideoGen-Bench作为专用评估套件,用于衡量动作质量和时间一致性。在多个TI2V后端和自演化策略下的实验显示,在ActVideoGen-Bench和VBench上均取得一致提升,证明了SPIRAL的有效性。
英文摘要
Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons, requiring procedural ordering, persistent action execution, and scene consistency beyond conventional TI2V's short-term fidelity. Existing single-shot video generation models typically operate in an open-loop manner, leading to incomplete action execution, hallucinated motions, and temporal drift. To address this, we propose SPIRAL, a closed-loop framework that performs sequential planning and iterative reflection for action-conditioned long-horizon video generation. Specifically, SPIRAL instantiates a think-act-reflect process: a PlanAgent decomposes high-level goals into sub-actions, which condition a VideoGenerator to synthesize each segment alongside a memory context, while a CriticAgent evaluates intermediate video segments to provide corrective feedback for iterative refinement. This closed-loop design further supports self-evolution by utilizing PlanAgent-proposed actions and CriticAgent-derived rewards for GRPO-based post-training to enhance the video generator's long-horizon consistency. Moreover, we introduce ActVideoGen-Dataset for task-specific training, and establish ActVideoGen-Bench as a dedicated evaluation suite for measuring action quality and temporal coherence. Experiments across multiple TI2V backbones alongside the self-evolving strategy show consistent gains on ActVideoGen-Bench and VBench, demonstrating the effectiveness of SPIRAL.