DeforM:通过时空掩码实现推理引导的物理感知视频生成
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对视频生成模型难以生成物理感知视频的问题,提出DeforM框架,引入VLM引导的物理推理模块及两种策略,经实验验证该框架能提高生成变形场景的真实感,优于基线模型。
AI中文摘要:
视频生成模型视觉质量高,但生成物理感知视频有困难。复杂变形动力学难以合成,缺乏对动态区域定位的物理推理会致生成失败。本文提出DeforM,一个推理引导的图像到视频生成框架,将模型焦点引向物理关键区域。引入VLM引导的物理推理模块DeforM-Reason识别目标对象并生成时空掩码。开发了DeforM-Free和DeforM-Injection两种策略。实验结果表明DeforM提高了生成变形场景的真实感,在视觉质量和物理一致性上优于基线模型。
英文摘要:
Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.