发表机构
Seoul National University; Sejong University(首尔大学; 世宗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频生成模型难以描绘物理交互的问题,提出利用合成状态转换框架,通过分类法和图像编辑模型创建起止状态锚点,并采用状态引导采样技术,提升模型生成合理交互的能力。
AI 中文摘要
虽然最近的视频生成模型能够合成高保真度的视频,但它们难以描绘合理的物理交互以及由此产生的状态转换,这是机器人和VR/AR应用中的一个关键瓶颈。为解决这一问题,我们引入了一个框架,用于生成可控交互的可扩展合成数据集。我们的流程利用结构化分类法和最先进的图像编辑模型来创建明确的“开始”和“结束”状态图像,这些图像作为交互的视觉锚点。为了利用这些锚点生成无缝视频,我们提出了状态引导采样(SGS),一种新颖的采样技术,可减轻朴素条件生成中常见的伪影。此外,我们开发并验证了一个新的自动化评估系统,该系统与人类判断一致,以确保数据质量。实验表明,在我们的数据集上微调基础模型显著增强了其生成合理交互的能力。
英文摘要
While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit `start' and `end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.
CommentsIJCAI 2026