发表机构
Zhejiang University; vivo Mobile Communication Co., Ltd.(浙江大学; 维沃移动通信有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频生成模型控制物理行为难的问题,提出VIPER框架,利用多模态大语言模型提取物理线索,通过分层训练引导图像到视频生成器,构建VIPER-19K数据集,实验证明该框架能实现物理行为转移且效果良好。
AI 中文摘要
现代视频生成模型能合成视觉上吸引人且时间连贯的片段,但用标准文本和图像条件控制其物理行为仍困难。核心挑战是条件瓶颈:材料响应等物理线索连续且相关,难用语言详尽指定却能由视频自然展示。我们提出VIPER,一种用于参考引导的图像到视频生成的视觉上下文物理推理框架。给定目标图像等,它将参考视为所需物理过程的密集视觉演示,用多模态大语言模型提取物理线索,通过分层训练策略引导预训练的图像到视频生成器,实现物理行为转移。我们构建了VIPER-19K数据集。实验表明VIPER比基线有更强的参考视频物理相似性和更高的人类偏好,定性结果也证明其能将参考的物理行为转移到新目标场景。
英文摘要
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.
Commentstech report