VideoCoCo:基于智能体双引擎系统的、以代码为思维链(CoT)的物理一致性视频生成方法
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
浏览论文内容
中文总结 AI 辅助
VideoCoCo作为智能体双引擎框架,以可执行Blender代码为过程级思维链,构建专属数据集提升视频编辑器适配性,在两个基准上显著优于基线,实现了高质量物理一致性视频生成。
中文摘要 AI 辅助
文本到视频模型已实现出色的视觉质量,但仍难以生成物理一致的动态效果,因为场景的时间演化需从高度压缩的文本提示中隐式推断。现有思维链方法引入中间计划或视觉状态,但这些表示通常不可执行或时间稀疏,限制了实例化和控制完整时空过程的能力。为解决此局限,我们提出VideoCoCo,这是一种智能体双引擎框架,其中可执行的Blender代码作为过程级思维链。给定文本提示后,编码智能体会合成一个Blender程序,明确指定场景及其时间演化。可执行的仿真引擎运行该程序以生成确定性时空草稿,随后生成式视频引擎通过草稿条件编辑将其转换为照片级真实感视频。这种分解将过程级推理与高保真视觉实现分离。为使视频编辑器适配仿真草稿,我们构建了VideoCoCo-3K,这是一个精心整理的草稿-指令-目标三元组数据集。VideoCoCo在PhyGenBench上将OmniWeaving基线从0.475提升至0.558,在VBench-2.0上从52.18提升至77.88,在两个基准上均取得最佳平均分数。这些结果表明,可执行代码为物理一致性视频生成提供了有效、可控且可检查的中间表示。
英文摘要
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
发表机构
- CUHK(香港中文大学)
- USTC(中国科学技术大学)
- SCUT(华南理工大学)
- HKU(香港大学)
- NTU(南洋理工大学)
- PKU(北京大学)
- SJTU(上海交通大学)
- CMU(卡内基梅隆大学)
- THU(清华大学)
- HUST(华中科技大学)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。