发表机构
Simon Fraser University; Stony Brook University; Lawrence Berkeley National Laboratory; Meta Reality Labs(西蒙菲莎大学; 石溪大学; 劳伦斯伯克利国家实验室; 元宇宙现实实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Fluid-Gen-Zero是一种无训练的即插即用框架,通过两级智能体工作流连接物理模拟器与预训练视频生成器,在流体-物体交互视频生成任务中显著提升模拟对齐度与人类偏好。
AI 中文摘要
我们提出了Fluid-Gen-Zero,这是一种用于物理感知流体-物体交互视频生成的无训练框架,它将物理推理与外观合成解耦。我们的核心见解是将运动动力学任务委托给物理模拟器,同时保留预训练视频生成器的外观建模能力。我们通过两级智能体工作流连接这两个领域:生成时规划,其中视觉语言模型(VLM)智能体解释意图和模拟滚动结果以组织生成片段;以及潜在空间引导,通过区域感知潜在封装将模拟信号注入去噪过程。这种即插即用设计与当前的视频基础模型兼容。我们还引入了一个用于流体-物体交互视频生成的基准。在基于CogVideoX的Tora、VACE和基于Wan的WanMove上,Fluid-Gen-Zero始终提高了模拟对齐度,将物体轨迹误差降低了26.7%-81.5%,流体fEPE(流体流端点误差)降低了67.9%-84.0%,同时在很大程度上保留了感知质量。在一项人类偏好研究中,在三种骨干网络的相同骨干网络比较中,评分者在55.1%-74.4%的案例中倾向于Fluid-Gen-Zero;在与基于模拟的方法的比较中,这一比例达到90.4%-94.2%。代码和数据将在论文接收后发布。
英文摘要
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.